<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="review-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">68024</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2025.068024</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Review</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>A Survey of Deep Learning for Time Series Forecasting: Theories, Datasets, and State-of-the-Art Techniques</article-title>
<alt-title alt-title-type="left-running-head">A Survey of Deep Learning for Time Series Forecasting: Theories, Datasets, and State-of-the-Art Techniques</alt-title>
<alt-title alt-title-type="right-running-head">A Survey of Deep Learning for Time Series Forecasting: Theories, Datasets, and State-of-the-Art Techniques</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Lu</surname><given-names>Gaoyong</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Ou</surname><given-names>Yang</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Wang</surname><given-names>Zhihong</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Qu</surname><given-names>Yingnan</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Xia</surname><given-names>Yingsheng</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-6" contrib-type="author">
<name name-style="western"><surname>Tang</surname><given-names>Dibin</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-7" contrib-type="author">
<name name-style="western"><surname>Kotenko</surname><given-names>Igor</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-8" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Li</surname><given-names>Wei</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><xref ref-type="aff" rid="aff-4">4</xref><email>wei.li@hrbeu.edu.cn</email></contrib>
<aff id="aff-1"><label>1</label><institution>The 10th Research Institute of China Electronics Technology Group</institution>, <addr-line>Chengdu, 610036</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>College of Computer Science and Technology, Harbin Engineering University</institution>, <addr-line>Harbin, 150001</addr-line>, <country>China</country></aff>
<aff id="aff-3"><label>3</label><institution>Laboratory of Computer Security Problems, St. Petersburg Federal Research Center of the Russian Academy of Sciences (SPC RAS)</institution>, <addr-line>Saint-Petersburg, 199178</addr-line>, <country>Russia</country></aff>
<aff id="aff-4"><label>4</label><institution>Modeling and Emulation in E-Government National Engineering Laboratory, Harbin Engineering University</institution>, <addr-line>Harbin, 150001</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Wei Li. Email: <email>wei.li@hrbeu.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2025</year></pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>23</day><month>09</month><year>2025</year></pub-date>
<volume>85</volume>
<issue>2</issue>
<fpage>2403</fpage>
<lpage>2441</lpage>
<history>
<date date-type="received">
<day>19</day>
<month>5</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>07</day>
<month>8</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_68024.pdf"></self-uri>
<abstract>
<p>Deep learning (DL) has revolutionized time series forecasting (TSF), surpassing traditional statistical methods (e.g., ARIMA) and machine learning techniques in modeling complex nonlinear dynamics and long-term dependencies prevalent in real-world temporal data. This comprehensive survey reviews state-of-the-art DL architectures for TSF, focusing on four core paradigms: (1) Convolutional Neural Networks (CNNs), adept at extracting localized temporal features; (2) Recurrent Neural Networks (RNNs) and their advanced variants (LSTM, GRU), designed for sequential dependency modeling; (3) Graph Neural Networks (GNNs), specialized for forecasting structured relational data with spatial-temporal dependencies; and (4) Transformer-based models, leveraging self-attention mechanisms to capture global temporal patterns efficiently. We provide a rigorous analysis of the theoretical underpinnings, recent algorithmic advancements (e.g., TCNs, attention mechanisms, hybrid architectures), and practical applications of each framework, supported by extensive benchmark datasets (e.g., ETT, traffic flow, financial indicators) and standardized evaluation metrics (MAE, MSE, RMSE). Critical challenges, including handling irregular sampling intervals, integrating domain knowledge for robustness, and managing computational complexity, are thoroughly discussed. Emerging research directions highlighted include diffusion models for uncertainty quantification, hybrid pipelines combining classical statistical and DL techniques for enhanced interpretability, quantile regression with Transformers for risk-aware forecasting, and optimizations for real-time deployment. This work serves as an essential reference, consolidating methodological innovations, empirical resources, and future trends to bridge the gap between theoretical research and practical implementation needs for researchers and practitioners in the field.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Time series forecasting</kwd>
<kwd>deep learning</kwd>
<kwd>transformer</kwd>
<kwd>neural network</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Natural Science Foundation of Heilongjiang Province</funding-source>
<award-id>LH2023F020</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Time series forecasting (TSF) stands as a pivotal analytical tool, enabling the prediction of future trends and patterns based on historical data. Its applications span a multitude of domains, including finance [<xref ref-type="bibr" rid="ref-1">1</xref>], economics [<xref ref-type="bibr" rid="ref-2">2</xref>], marketing [<xref ref-type="bibr" rid="ref-3">3</xref>], social sciences [<xref ref-type="bibr" rid="ref-4">4</xref>] and environmental sciences [<xref ref-type="bibr" rid="ref-5">5</xref>]. In marketing, it aids in forecasting sales, market shares, and advertising effects. Moreover, TSF is crucial for predicting meteorological events such as rainfall [<xref ref-type="bibr" rid="ref-6">6</xref>], traffic flow [<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-8">8</xref>], medical drug response, and the operational demands of diverse systems, underscoring its versatility and critical role in both scientific research and practical applications.</p>
<p>The emergence of deep learning (DL) has significantly advanced the methodologies and efficacy of TSF. DL, a prominent branch of machine learning, seeks to simulate and understand the human brain&#x2019;s functioning by constructing and training multi-layer neural networks [<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-10">10</xref>]. These networks depict complex relationships in data through connections and weights between layers, learning these relationships through training on large-scale data. Its ability to identify patterns and features in large datasets enables highly accurate predictions and classifications, automatically extracting and representing useful information.</p>
<p>The methodological landscape of TSF has undergone significant transformation alongside the exponential growth of temporal data and its increasing dimensionality. Early approaches predominantly employed conventional statistical techniques that incorporated numerous assumptions about data patterns, frequently proving inadequate for practical applications because they failed to model nonlinear dynamics effectively. Subsequent advancements introduced machine learning (ML) paradigms, including Support Vector Machines (SVMs) [<xref ref-type="bibr" rid="ref-11">11</xref>], Gradient Boosted Regression Trees (GBRT), and Hidden Markov Models (HMMs) [<xref ref-type="bibr" rid="ref-12">12</xref>], which demonstrated improved performance. Nevertheless, these ML methods still exhibited constraints when processing intricate temporal structures and extended dependencies characteristic of real-world time series.</p>
<p>With the advent of DL, a new era in TSF began. DL techniques, particularly those used in natural language processing [<xref ref-type="bibr" rid="ref-13">13</xref>], have been effectively applied to time series research, significantly improving the nonlinear modeling capabilities of TSF methods. The development of DL-based TSF algorithms is shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. These advancements have made DL an effective solution for solving TSF problems, enabling more accurate and reliable predictions. Despite extensive research, comprehensive and reviews summarizing the overall progress and state-of-the-art techniques in TSF remain scarce. This paper addresses this gap by providing an exhaustive and survey of TSF based on deep learning techniques. We introduce foundational concepts and definitions, classify TSF methods according to different models, and review state-of-the-art (SOTA) methods. Additionally, we discuss frequently used datasets and performance evaluation metrics across various domains, offering a comprehensive overview of the field.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>The development of DL-based TSF algorithms</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_68024-fig-1.tif"/>
</fig>
<p>Compared to other surveys [<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-14">14</xref>,<xref ref-type="bibr" rid="ref-15">15</xref>], we offer a more comprehensive and accessible overview, integrating specific application scenarios and placing particular emphasis on recent advanced topics. The contributions of this paper are as follows:
<list list-type="bullet">
<list-item>
<p><bold>Overview of Time Series Forecasting Tasks:</bold> We provide clear definitions of time series data and TSF, introducing and explaining the TSF task from various perspectives to enhance understanding.</p></list-item>
<list-item>
<p><bold>New Classification and Comprehensive Review:</bold> We propose a new classification for TSF research, categorizing methods into CNN-based models, RNN-based models, GNN-based models, Transformer-based models, composite models, and other forecasting models. We summarize key issues in TSF and discuss relevant work and effective solutions for each.</p></list-item>
<list-item>
<p><bold>Comprehensive Resources:</bold> We gather extensive resources, providing detailed introductions and reviews on various aspects of TSF, including definitions, commonly used models, and evaluation metrics. We also present the SOTA models from recent years, enabling scholars to gain a deeper understanding of TSF tasks and apply their knowledge in practice.</p></list-item>
<list-item>
<p><bold>Future Research Directions:</bold> Based on our research and experimental findings, we outline potential future research directions, providing a forward-looking perspective on the field.</p></list-item>
</list></p>
<p>The remaining structure of the article is as follows: In <xref ref-type="sec" rid="s2">Section 2</xref>, we define time series data and TSF, explaining the TSF task from various perspectives. In <xref ref-type="sec" rid="s3">Section 3</xref>, we classify TSF methods according to different forecasting models and review state-of-the-art TSF methods [<xref ref-type="bibr" rid="ref-16">16</xref>]. In <xref ref-type="sec" rid="s4">Section 4</xref>, we discuss and evaluate the datasets frequently used in TSF across various domains, along with their performance evaluation metrics. In <xref ref-type="sec" rid="s5">Section 5</xref>, we speculate on the future trends of TSF research and briefly summarize the core findings of this study.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Problem Definition</title>
<p>In this section, we introduce the foundational concepts and definitions of the TSF problem. Specifically, we define data and its characteristics in <xref ref-type="sec" rid="s2_1">Section 2.1</xref>, and discuss the definition of TSF and various approaches to the TSF task based on different classification criteria in <xref ref-type="sec" rid="s2_2">Section 2.2</xref>.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Time Series Definition</title>
<p>Time series analysis serves as a fundamental methodology across diverse disciplines [<xref ref-type="bibr" rid="ref-16">16</xref>], ranging from quantitative finance and meteorological science to biomedical engineering and industrial applications. These temporally ordered observations record the evolution of system states, enabling researchers to identify underlying patterns, detect anomalies, and predict future behaviors. The analytical value of time series stems from their ability to represent dynamic processes through sequential measurements.
<list list-type="bullet">
<list-item>
<p><bold>Trend:</bold> The trend component in time series analysis characterizes the persistent [<xref ref-type="bibr" rid="ref-16">16</xref>], long-term movement of data values, revealing either growth, decline, or stationary behavior across extended observation periods. This fundamental feature serves as a critical indicator of changes in the observed system, offering researchers meaningful information about the intrinsic dynamics of temporal processes.</p>
<p>For example, in equity market analysis, securities frequently demonstrate prolonged bullish trends during economic expansions, contrasted by bearish trends during contractions. These directional movements encapsulate the fundamental shifts in market valuation. As shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>, our analysis of product sales data reveals a consistent growth pattern, where the fitted trend line (dotted) clearly tracks the upward trajectory of actual sales measurements (solid line) over multiple business cycles.</p>

<p>Trends can manifest as either local or global patterns, and a single time series can exhibit both. For example, in the sales volume data, the overall trend is upward, but there may be localized upward trends during peak seasons and localized downward trends during off-seasons. Additionally, trends can be linear or non-linear [<xref ref-type="bibr" rid="ref-16">16</xref>]. A linear trend is characterized by a constant rate of increase or decrease, while a non-linear trend typically exhibits a multiplicative increase, meaning the rate of change is proportional to the previous values. Understanding the nature of the trend, whether it is linear or non-linear, is crucial for selecting appropriate forecasting models and strategies.</p></list-item>
<list-item>
<p><bold>Seasonality:</bold> Seasonality in time-series data refers to the cyclical repeating patterns or periodic changes that occur over specific time frames. These cyclical variations are typically associated with particular seasons, months, days of the week, or other temporal units. For instance, in temperate regions, higher temperatures during the summer and lower temperatures during the winter create a distinct seasonal pattern. This pattern is crucial for weather forecasting and agricultural planning, among other applications. Similarly, stock markets are influenced by seasonality. The release of quarterly reports, annual reports, and financial data can significantly impact stock prices and trading volumes. Additionally, during specific seasons, such as the end of the year or the beginning of the year, investors may adjust their portfolios to accommodate seasonal changes. <xref ref-type="fig" rid="fig-3">Fig. 3</xref> illustrates the seasonal cycle of a product&#x2019;s sales volume, highlighting the regular fluctuations that occur within a specific time frame.</p>
</list-item>
<list-item>
<p><bold>Randomness:</bold> The randomness of a time series refers to the absence of a clear pattern, trend, or periodicity in the data over time. Instead, the data exhibit a high degree of variability and unpredictability, characterized by random fluctuations and irregularities. This randomness implies that the movements of the data are not governed by fixed patterns or laws, making it difficult to predict future values based solely on historical data. Randomness can be caused by a variety of factors, including external shocks, measurement errors, or inherent stochastic processes.</p></list-item>
</list><fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Trend in sales volume of a product</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_68024-fig-2.tif"/>
</fig><fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Seasonality in sales volume of a product</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_68024-fig-3.tif"/>
</fig></p>
<p>From a mathematical perspective, a time series is a sequence of random variables, either finite or infinite, arranged in chronological order. These variables represent the values of a specific statistical indicator over time, typically sampled at a relatively fixed frequency to capture the process of change. Mathematically, a time series can be conceptualized as a stochastic process, where each observation is a random variable influenced by various factors.</p>
<p>The time series data can be represented in matrix form as follows. When there is only one sensor, the time series is one-dimensional:<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the value of the sensor at the current timestamp, and <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents a set of time series. This can be extended to multidimensional time series:<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>:</mml:mo><mml:mi>N</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>n</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>n</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>n</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>n</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></disp-formula>where <italic>N</italic> denotes an <italic>N</italic>-dimensional time series, <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>n</mml:mi></mml:math></inline-formula> represents a time series with <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>n</mml:mi></mml:math></inline-formula> sample data, and <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>t</mml:mi></mml:math></inline-formula> denotes the current time step. <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes a vector, <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>n</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>n</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>, and <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the time series point of the time series at time <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>t</mml:mi></mml:math></inline-formula>. <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>:</mml:mo><mml:mi>N</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>n</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes <italic>N</italic>-dimensional time series data consisting of <italic>N</italic> time series.</p>
<p>Time series data can be either one-dimensional or multidimensional. A one-dimensional time series, focusing on a single variable, is represented as:<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the time series of the past <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>k</mml:mi></mml:math></inline-formula> time series. This representation is particularly useful for analyzing the behavior of a single variable over time, such as the daily closing price of a stock or the hourly temperature readings at a specific location. Understanding the mathematical structure of time series data is fundamental for developing effective models and techniques for time series forecasting.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Definition of Time Series Forecasting</title>
<p>TSF serves as a fundamental analytical technique for temporal data analysis, with wide-ranging applications including classification tasks, anomaly identification, and future value prediction. At its core, TSF methodologies employ historical observations to model temporal patterns, enabling the projection of future system states through careful analysis of established trends and behavioral characteristics. The primary objective involves generating accurate predictions for time steps beyond a given reference point <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>t</mml:mi></mml:math></inline-formula>, utilizing only information available prior to <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi>t</mml:mi></mml:math></inline-formula>.</p>
<p>For clarity of exposition, we focus our discussion on single-step forecasting scenarios. The underlying mathematical formulation can be expressed as:<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>S</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi>i</mml:mi></mml:math></inline-formula> denotes the <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>i</mml:mi></mml:math></inline-formula>-th set of time series data, <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> represents the forecasted value at time <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> based on the data from <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi></mml:math></inline-formula> to <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mi>t</mml:mi></mml:math></inline-formula>, <italic>S</italic> represents the static factors that remain constant during the forecasting process, and <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the prediction function of the model. Since the static factor <italic>S</italic> remains unchanged during the forecasting process, in order to simplify the formulas for better understanding, <italic>S</italic> will be omitted in the subsequent equations.</p>
<p>TSF methodologies can be classified according to two primary dimensions: (1) prediction horizon, encompassing single-step (immediate next point) and multi-step (extended future sequence) forecasting approaches, and (2) input variable configuration, distinguishing between autoregressive (target series only) and covariate-based (incorporating external factors) methods. This categorization reflects fundamental differences in problem formulation and technical requirements, with each paradigm presenting unique computational challenges and application scenarios that will be explored in subsequent sections.</p>
<sec id="s2_2_1">
<label>2.2.1</label>
<title>Single-Step Forecasting and Multi-Step Forecasting</title>
<p>TSF methodologies can be fundamentally classified by their prediction horizon, with single-step forecasting (one-step-ahead prediction) focusing on immediate future values and multi-step forecasting (also called long-horizon prediction) addressing sequences of future observations. Each type of forecasting serves different purposes and is suited to different scenarios.</p>
<p><bold>(1) Single-Step Forecasting</bold></p>
<p>Single-step forecasting focuses on estimating the immediate subsequent value in a time series using historical observations (<xref ref-type="fig" rid="fig-4">Fig. 4</xref>). This methodology is particularly valuable in applications requiring high temporal resolution, including financial market analysis and meteorological prediction systems, where accurate near-term projections are essential for operational decision-making. The technique&#x2019;s emphasis on the most proximate future point enables enhanced prediction accuracy for real-time applications, as it avoids the compounding errors associated with longer forecasting horizons while providing the low-latency outputs needed for time-sensitive scenarios.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Single-step forecasting</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_68024-fig-4.tif"/>
</fig>
<p>The mathematical expression for single-step forecasting is as follows:<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>S</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi>i</mml:mi></mml:math></inline-formula> denotes the number of input features, <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mi>j</mml:mi></mml:math></inline-formula> is the number of output features, <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mi>k</mml:mi></mml:math></inline-formula> is the time step parameter [<xref ref-type="bibr" rid="ref-9">9</xref>], and <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mi>t</mml:mi></mml:math></inline-formula> is the current time step, <italic>S</italic> denotes any static factors that remain constant during the forecasting process.</p>
<p><bold>(2) Multi-Step Forecasting</bold></p>
<p>Multi-step forecasting extends predictive capabilities by generating sequential future values from historical observations (<xref ref-type="fig" rid="fig-5">Fig. 5</xref>), addressing scenarios requiring extended-horizon projections. This methodology is indispensable for applications where capturing evolving trends is critical, such as electricity load planning, macroeconomic policy formulation, and investment portfolio optimization. By simultaneously modeling multiple future states, it enables proactive resource allocation and risk mitigation, though it introduces unique challenges like error accumulation and temporal dependency decay that require specialized architectural solutions.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Multi-step forecasting</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_68024-fig-5.tif"/>
</fig>
<p>The mathematical expression for multi-step forecasting is as follows:<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>:</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>S</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi>s</mml:mi></mml:math></inline-formula> denotes the prediction horizon, representing the number of future time steps to be predicted, <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>:</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> represents the predicted values for the <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mi>j</mml:mi></mml:math></inline-formula>-th output feature from time <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> to <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>.</p>
<p>Multi-step forecasting can be implemented using several methods, each with its own advantages and disadvantages. The four main methods are:
<list list-type="order">
<list-item>
<p><bold>Direct Multi-Step Forecasting:</bold> This method involves building a separate model for each future time step. Each model is trained to predict a specific time step, and the predictions are aggregated to form the final forecast. While this approach can capture the unique patterns of each time step, it can become computationally intensive when the prediction horizon is large, as it requires training multiple models.</p></list-item>
<list-item>
<p><bold>Recursive Multi-Step Forecasting:</bold> In this method, a single model is used to predict the next time step, and the predicted value is then fed back into the model as an input for the subsequent prediction. This process is repeated recursively to generate predictions for multiple future time steps. While this approach is computationally efficient, it can accumulate errors over time, leading to less accurate long-term predictions.</p></list-item>
<list-item>
<p><bold>Direct-Recursive Hybrid Multi-Step Forecasting:</bold> This method combines the direct and recursive approaches. It uses a separate model for each time step but also incorporates the predictions from previous time steps as inputs [<xref ref-type="bibr" rid="ref-10">10</xref>]. This hybrid approach aims to balance the computational efficiency of the recursive method with the accuracy of the direct method.</p></list-item>
<list-item>
<p><bold>Multi-Output Forecasting:</bold> This method involves training a single model to predict all future time steps simultaneously. The model is designed to handle multiple outputs, capturing the relationships between different time steps. This approach can be more efficient and accurate, especially when the relationships between time steps are strong.</p></list-item>
</list></p>
<p>These methodologies exhibit distinct advantages and limitations, with optimal model selection contingent upon both dataset properties (e.g., temporal resolution, noise characteristics) and application requirements (e.g., prediction horizon, interpretability needs).</p>
<p><bold>Direct Multi-Step Forecasting.</bold> To build a separate model for each future time step, direct multi-step forecasting involves training each model to predict a specific time step. Essentially, this method is a series of single-step forecasts, with each model trained on the same historical data but predicting a different future time step, as shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>. The mathematical expression for direct multi-step forecasting is as follows.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Direct multi-step forecasting</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_68024-fig-6.tif"/>
</fig>
<p>The mathematical expression for direct multi-step prediction is as follows:<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mi></mml:mi><mml:mspace width="2em" /><mml:mo>&#x2026;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mi>i</mml:mi></mml:math></inline-formula> denotes the number of input features, <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mi>k</mml:mi></mml:math></inline-formula> denotes the time step and <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mi>t</mml:mi></mml:math></inline-formula> denotes the current time step, and <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mi>s</mml:mi></mml:math></inline-formula> denotes the predicted horizon.</p>
<p>While direct multi-step forecasting can capture the unique patterns of each time step, it can become computationally intensive when the prediction horizon is large, as it requires training multiple models. Additionally, since each model is trained independently, it may not capture the correlations between different time steps, potentially leading to less accurate predictions.</p>
<p><bold>Recursive Multi-Step Forecasting.</bold> It employs an iterative prediction strategy where a single-step model generates successive forecasts by recursively feeding its outputs as inputs for subsequent time steps (<xref ref-type="fig" rid="fig-7">Fig. 7</xref>). This chained prediction mechanism can be formalized as:</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Recursive multi-step forecasting</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_68024-fig-7.tif"/>
</fig>
<p>The mathematical expression is as follows:<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mi></mml:mi><mml:mspace width="2em" /><mml:mo>&#x2026;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mi>i</mml:mi></mml:math></inline-formula> denotes the number of input features, <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mi>k</mml:mi></mml:math></inline-formula> denotes the time step and <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mi>t</mml:mi></mml:math></inline-formula> denotes the current time step, and <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mi>s</mml:mi></mml:math></inline-formula> denotes the predicted horizon [<xref ref-type="bibr" rid="ref-9">9</xref>].</p>
<p>While recursive multi-step forecasting is computationally efficient, it can accumulate errors over time, leading to less accurate long-term predictions. This is because the model uses predicted values rather than actual observations as inputs for subsequent predictions.</p>
<p><bold>Direct-Recursive Hybrid Multi-Step Forecasting. </bold>It approach enhances multi-step forecasting accuracy by synergistically integrating the strengths of both direct and recursive methodologies. This framework employs a cascade of specialized models, where each subsequent predictor incorporates outputs from preceding models as supplementary inputs (<xref ref-type="fig" rid="fig-8">Fig. 8</xref>). The mathematical formulation of this hybrid paradigm can be expressed as:</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Direct-recursive hybrid multi-step forecasting</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_68024-fig-8.tif"/>
</fig>
<p>The mathematical expression for direct-recursive hybrid multi-step prediction [<xref ref-type="bibr" rid="ref-10">10</xref>] is as follows:<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>:</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mi></mml:mi><mml:mspace width="2em" /><mml:mo>&#x2026;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mi>s</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mi>i</mml:mi></mml:math></inline-formula> denotes the number of input features, <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mi>k</mml:mi></mml:math></inline-formula> denotes the time step and <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mi>t</mml:mi></mml:math></inline-formula> denotes the current time step, and <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mi>s</mml:mi></mml:math></inline-formula> denotes the predicted horizon.</p>
<p>This hybrid approach aims to balance the computational efficiency of the recursive method with the accuracy of the direct method by leveraging the strengths of both approaches.</p>
<p><bold>Multi-Output Multi-Step Forecasting.</bold> Training a single model to predict all future time steps simultaneously, multi-output multi-step forecasting is particularly useful in neural network models can handle multidimensional inputs and outputs. Computationally efficient, this method captures the relationships between different time steps, making it a powerful tool for long-term forecasting.</p>
</sec>
<sec id="s2_2_2">
<label>2.2.2</label>
<title>Autoregressive Forecasting and Covariate Forecasting</title>
<p>TSF methodologies can be fundamentally classified by their input variable structure into two paradigms: (1) autoregressive approaches that exclusively utilize historical values of the target series, and (2) covariate-based methods that incorporate external explanatory variables. The selection between these approaches requires careful consideration of multiple factors, including data dimensionality, feature availability, and the required prediction horizon, as each technique exhibits distinct advantages in handling different temporal patterns and application scenarios.</p>
<p><bold>(1) Autoregressive forecasting</bold></p>
<p>Autoregressive forecasting involves using only the time series data itself for prediction, without considering other external factors, as illustrated in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>. By leveraging historical data, an autoregressive model can be built to predict future values. These models typically use linear regression methods to fit the historical data and make predictions based on the fitting results.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Autoregressive forecasting</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_68024-fig-9.tif"/>
</fig>
<p>The advantages of autoregressive forecasting include its simplicity, ease of use, and computational efficiency. However, it has some limitations. It can capture linear relationships and is well-suited for handling nonlinear data. Additionally, it cannot effectively model seasonal or trend components. Therefore, when using autoregressive forecasting, it is important to select appropriate models and parameters based on the specific characteristics of the data. Autoregressive forecasting is typically used for short-term predictions, as it focuses solely on the patterns within the time series data and does not account for external influences.</p>
<p>The mathematical expression for autoregressive prediction is as follows:<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mi>t</mml:mi></mml:math></inline-formula> denotes the current time step and <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mi>k</mml:mi></mml:math></inline-formula> denotes the input time step size [<xref ref-type="bibr" rid="ref-9">9</xref>]. <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:mi>f</mml:mi></mml:math></inline-formula> represents the algorithm of the model in the current application scenario.</p>
<p><bold>(2) Covariate forecasting</bold></p>
<p>Covariate forecasting extends the forecasting model by incorporating additional external factors, or covariates, in addition to the time series data itself, as shown in <xref ref-type="fig" rid="fig-10">Fig. 10</xref>. These covariates can include variables. By modeling the relationships between these covariates and the target variable, covariate forecasting can provide more accurate and comprehensive predictions.</p>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Covariate forecasting</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_68024-fig-10.tif"/>
</fig>
<p>The primary advantage of covariate forecasting is its ability to improve prediction accuracy by considering a broader range of relevant factors. For example, in sales forecasting, covariates such as weather, economic indices, and promotional activities can be included to enhance the predictive power of the model. It requires the collection and processing of additional data, which can increase the complexity and computational cost. Moreover, selecting the appropriate covariates is crucial and often requires domain expertise and careful analysis.</p>
<p>The mathematical expression for covariate prediction is given below:<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mi>i</mml:mi></mml:math></inline-formula> indicates the number of input features, <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mi>k</mml:mi></mml:math></inline-formula> means the input time step size [<xref ref-type="bibr" rid="ref-9">9</xref>], and <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mi>t</mml:mi></mml:math></inline-formula> refers to the current time step.</p>
<p>In summary, autoregressive forecasting is a simple and efficient method suitable for short-term predictions when the focus is on the inherent patterns of the time series data. Covariate forecasting, on the other hand, is more complex but can provide higher accuracy by incorporating external factors that influence the target variable.</p>
</sec>
<sec id="s2_2_3">
<label>2.2.3</label>
<title>Isometric Interval Forecasting and Non-Isometric Interval Forecasting</title>
<p>Time series forecasting (TSF) tasks can also be classified based on the equality of the time intervals between observations in the time series data. The two main categories are isometric interval forecasting and non-isometric interval forecasting, which are described below.</p>
<p><bold>(1) Isometric Interval Forecasting</bold></p>
<p>Isometric interval forecasting involves predicting future values based on time series data where the time interval between each observation is equal, as illustrated in <xref ref-type="fig" rid="fig-11">Fig. 11</xref>. This method is particularly suitable for data that exhibit periodic patterns. When the data have significant periodic variations, isometric interval forecasting can effectively capture these patterns and use them to predict future values.</p>
<fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>Isometric interval forecasting</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_68024-fig-11.tif"/>
</fig>
<p>The primary advantage of isometric interval forecasting is its ability to capture and utilize periodic patterns in the data. For instance, in the field of finance, stock prices often exhibit cyclical behavior, and isometric interval forecasting can be effectively used to predict future stock prices. Similarly, in meteorology, temperature and rainfall data often show cyclical variations, making isometric interval forecasting a suitable method for predicting future weather conditions.</p>
<p>The mathematical expression for isometric interval forecasting is as follows:<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mi>i</mml:mi></mml:math></inline-formula> is the <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mi>i</mml:mi></mml:math></inline-formula>-th time series data, <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mi>t</mml:mi></mml:math></inline-formula> denotes the current time step, and <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mi>k</mml:mi></mml:math></inline-formula> is the input time step size.</p>
<p><bold>(2) Non-Isometric Interval Forecasting</bold></p>
<p>Non-isometric interval forecasting deals with time series data where the time interval between observations is unequal, as shown in <xref ref-type="fig" rid="fig-12">Fig. 12</xref>. This method is suitable for data that do not exhibit a clear periodic pattern. Non-isometric interval forecasting can be used to make predictions even when there is no apparent periodic variation in the data. For example, in the finance domain, stock trading data may be non-equally spaced.</p>
<fig id="fig-12">
<label>Figure 12</label>
<caption>
<title>Nonisometric interval forecasting</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_68024-fig-12.tif"/>
</fig>
<p>The advantage of non-isometric interval forecasting is its ability to handle non-equally spaced data and make predictions even in the absence of significant cyclical variations. However, this method has some limitations. First, it requires more complex techniques, such as interpolation, to handle the unequal intervals, which can increase computational complexity and resource requirements. Additionally, non-isometric interval forecasting may be less accurate than isometric interval forecasting because it cannot utilize periodic patterns in the data for prediction.</p>
<p>The mathematical expression for non-isometric interval forecasting is as follows:<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mi>i</mml:mi></mml:math></inline-formula> denotes the <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mi>i</mml:mi></mml:math></inline-formula>-th time series data, <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mi>t</mml:mi></mml:math></inline-formula> represents the current time step, <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mi>k</mml:mi></mml:math></inline-formula> is the input time step size, and <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mrow><mml:mo>{</mml:mo><mml:mi>l</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>k</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">N</mml:mi></mml:mrow><mml:mrow><mml:mo>+</mml:mo></mml:mrow></mml:msup><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>. Equalities may exist between some values of <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mi>l</mml:mi></mml:math></inline-formula>, but not all values before <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mi>l</mml:mi></mml:math></inline-formula> are identical, as there are varying sizes of <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mi>l</mml:mi></mml:math></inline-formula>.</p>
<p>In summary, isometric interval forecasting is suitable for data with regular and periodic patterns, while non-isometric interval forecasting is used for data with irregular or non-periodic patterns.</p>
</sec>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>A Review and Classification of DL Techniques in TSF</title>
<p>In this section, we present and discuss typical time series forecasting (TSF) approaches based on different deep learning models. We categorize TSF methods into five types: CNN-based methods, RNN-based methods, MLP-based methods, GNN-based methods, and Transformer-based methods. For each category, we explain the principles and roles of the methods in addressing TSF problems and provide a review of specific methods.
<list list-type="bullet">
<list-item>
<p><bold>CNN-Based Methods:</bold> Convolutional Neural Networks are designed to extract local features from time series data through convolutional layers. This method is particularly effective for capturing short-term dependencies in the data, making it suitable for tasks like financial prediction, where local patterns in price movements are important.</p></list-item>
<list-item>
<p><bold>RNN-Based Methods:</bold> Recurrent Neural Network is ideal for time series forecasting, because they can model temporal dependencies by maintaining a memory of past inputs.</p></list-item>
<list-item>
<p><bold>MLP-Based Methods:</bold> Multi-Layer Perceptrons are feedforward neural networks that can model both linear and nonlinear relationships in time series data. They are useful for forecasting tasks like sales prediction, where multiple input features, such as seasonality or promotional events, need to be combined to predict future outcomes.</p></list-item>
<list-item>
<p><bold>GNN-Based Methods:</bold> Graph Neural Network is designed to handle time series data with complex spatial-temporal dependencies. They are especially useful for tasks like traffic flow prediction.</p></list-item>
<list-item>
<p><bold>Transformer-Based Methods:</bold> Transformer models use self-attention mechanisms to handle long-term dependencies in time series data. Unlike RNNs, transformers can process sequences in parallel and efficiently model long-range dependencies.</p></list-item>
</list></p>
<p>Shallow networks are typically suitable for handling simple forecasting problems, with the advantage of lower computational cost and faster training speed. However, as the complexity of the problem increases, shallow networks may fail to capture the deeper relationships and patterns in the data. In such cases, deep networks become more necessary. Although deep networks have significantly higher computational costs than shallow networks, they can effectively learn more complex feature representations, especially when dealing with nonlinear, high-dimensional, or long-term dependent data. By introducing more layers, deep networks can capture more detailed patterns, and in many practical problems, deep networks have been shown to significantly improve forecasting accuracy.</p>
<p>The training and inference processes of DL models typically require substantial computational resources, especially when handling large datasets. The costs during training mainly stem from the model&#x2019;s complexity, the number of parameters, and the iterative processing of data. DLg models often require large memory and processing power. While inference is relatively lighter, highly complex models still consume significant time, particularly during real-time or large batch predictions. Additionally, using hardware accelerators like GPUs or TPUs during training incurs high energy consumption and operational costs. For simple problems, shallow networks remain an efficient choice, while for complex problems, deep networks offer stronger modeling capabilities. We not only conduct a detailed analysis and explanation of the models but also provide a side-by-side performance summary. We implement the model with the PyTorch toolkit on a Linux server with a GeForce RTX 4090 GPU.</p>
<sec id="s3_1">
<label>3.1</label>
<title>CNN-Based Methods</title>
<p>The structure of the Convolutional Neural Networks (CNNs) used for TSF is shown in <xref ref-type="fig" rid="fig-13">Fig. 13</xref>. For ease of understanding, a single-channel solution, as shown in <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>, is used.</p>
<fig id="fig-13">
<label>Figure 13</label>
<caption>
<title>The structure of the CNN used for the TSF task</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_68024-fig-13.tif"/>
</fig>
<p>The working principle of CNNs is based on convolutional operations and feature extraction. In TSF tasks, CNNs can capture patterns and local dependencies at different time scales and extract the most predictive features from time series data. Additionally, CNNs can handle multivariate time series data by using 2D CNNs, which can extract important feature patterns in both time and other dimensions simultaneously.</p>
<p>Temporal Convolutional Networks (TCNs) [<xref ref-type="bibr" rid="ref-17">17</xref>] are an improvement over traditional CNNs. TCNs use convolutional layers to capture patterns and dependencies in time series data and have shown strong performance in various time series tasks. The key improvements in TCNs include three blocks, as shown in <xref ref-type="fig" rid="fig-14">Figs. 14</xref>&#x2013;<xref ref-type="fig" rid="fig-16">16</xref>.</p>
<fig id="fig-14">
<label>Figure 14</label>
<caption>
<title>The structure of the causal convolution in TCN</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_68024-fig-14.tif"/>
</fig><fig id="fig-15">
<label>Figure 15</label>
<caption>
<title>The structure of the dilated convolution in TCN</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_68024-fig-15.tif"/>
</fig><fig id="fig-16">
<label>Figure 16</label>
<caption>
<title>The structure of the residual block in TCN</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_68024-fig-16.tif"/>
</fig>
<p><list list-type="bullet">
<list-item>
<p><bold>Causal Convolution:</bold> In causal convolution, the convolution kernel only slides over current and past time steps, ensuring that the model uses only past information to predict the current output, following a causal relationship.</p></list-item>
<list-item>
<p><bold>Dilated Convolution:</bold> Dilated convolution architectures enhance temporal dependency modeling through strategically spaced convolutional kernels, expanding the receptive field while maintaining parameter efficiency.</p></list-item>
<list-item>
<p><bold>Residual Blocks:</bold> Residual blocks use residual connections to solve the problem of vanishing gradients in deep networks. They allow gradients to flow directly to earlier layers, improving network performance and supporting the construction of deeper networks.</p></list-item>
</list></p>
<p>As shown in <xref ref-type="fig" rid="fig-15">Fig. 15</xref>, dilated convolution in Temporal Convolutional Networks is a technique that expands the receptive field by introducing gaps between the data points the kernel processes. This is done by applying a kernel with a specified &#x201C;dilation rate,&#x201D; allowing the model to capture long-range dependencies without increasing the computational load. By stacking dilated convolutions with increasing dilation rates, TCNs can model dependencies at various time scales. In the network at each update step, dropout forces the network to rely on multiple paths to make predictions. This helps to reduce the model&#x2019;s dependency on specific neurons, thereby improving generalization to unseen data. Dropout is typically applied in fully connected layers and can be controlled, which specifies the probability of deactivating a neuron. By introducing dropout, the network is encouraged to learn more robust features, especially in complex tasks with limited training data.</p>
<p>Some TSF works based on CNN and TCN modeling are discussed below, as shown in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>CNN-based model</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>Year</th>
<th>Type</th>
<th>Baseline</th>
</tr>
</thead>
<tbody>
<tr>
<td>DLSF</td>
<td>2017</td>
<td>Multivariate</td>
<td>Time series analyzing methods, SVM</td>
</tr>
<tr>
<td>Borovykh et al.</td>
<td>2017</td>
<td>Multivariate</td>
<td>LSTM, VAR</td>
</tr>
<tr>
<td>Dong et al.</td>
<td>2017</td>
<td>Multivariate</td>
<td>Linear Regression, SVR, NN, CNN</td>
</tr>
<tr>
<td>DSANet</td>
<td>2019</td>
<td>Multivariate</td>
<td>VAR, LRidge, LSVR, GP, GRU, LSTNet, TPA</td>
</tr>
<tr>
<td>MLCNN</td>
<td>2020</td>
<td>Multivariate</td>
<td>VAR, MTCNN, LSTNet, RNN-LSTM, AECRNN</td>
</tr>
<tr>
<td>SCINet</td>
<td>2022</td>
<td>Multivariate</td>
<td>Autoformer, Informer, Transformer, TCN, LSTNet, TPA-LSTM,<break/>Pyraformer,LogTrans, Reformer, ARIMA, Prophet, DeepAR, N-Beats</td>
</tr>
<tr>
<td>MICN</td>
<td>2023</td>
<td>Multivariate</td>
<td>FEDformer, Informer, LSTNet, LSTM, TCN, Autoformer, LogTrans</td>
</tr>
<tr>
<td>TimesNet</td>
<td>2023</td>
<td>Multivariate</td>
<td>LSTM, LSTNet, LSSL, TCN, LightTS, DLinear, Reformer, Informer,<break/>Pyraformer, Autoformer, FEDformer, Non-stationgary Transformer,<break/>ETSformer</td>
</tr>
<tr>
<td>LightCTS</td>
<td>2023</td>
<td>Multivariate</td>
<td>DCRNN, GWNet, AGCRN, MTGNN, AutoCTS, EnhanceNet, FOGS</td>
</tr>
<tr>
<td>Cross-LKTCN</td>
<td>2023</td>
<td>Multivariate</td>
<td>PatchTST, Dlinear, Crossformer, MTGNN, MICN, SCINet,<break/>FEDfomer, Autoformer</td>
</tr>
<tr>
<td>PatchMixer</td>
<td>2023</td>
<td>Univariate</td>
<td>PatchTST, DLinear, MICN, TimesNet, FEDformer,<break/>Autoformer, Informer</td>
</tr>
<tr>
<td>ModernTCN</td>
<td>2024</td>
<td>Multivariate</td>
<td>PatchTST, Crossformer, FEDformer, MTS-Mixer, LightTS, DLinear,<break/>RMLP, RLinear, TimesNet, MICN, SCINet</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>As shown in <xref ref-type="table" rid="table-2">Table 2</xref>, we evaluate the overall performance of the model on the ETTh1 dataset with a prediction length of 720.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Performance on the ETTh1 dataset. The prediction lengths are 720</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>MAE</th>
<th>MSE</th>
</tr>
</thead>
<tbody>
<tr>
<td>DLSF</td>
<td>1.701</td>
<td>1.699</td>
</tr>
<tr>
<td>Borovykh et al.</td>
<td>1.656</td>
<td>1.677</td>
</tr>
<tr>
<td>Dong et al.</td>
<td>1.267</td>
<td>1.197</td>
</tr>
<tr>
<td>DSANet</td>
<td>0.940</td>
<td>0.945</td>
</tr>
<tr>
<td>MLCNN</td>
<td>0.991</td>
<td>0.912</td>
</tr>
<tr>
<td>SCINet</td>
<td>0.527</td>
<td>0.544</td>
</tr>
<tr>
<td>MICN</td>
<td>0.491</td>
<td>0.499</td>
</tr>
<tr>
<td>TimesNet</td>
<td>0.450</td>
<td>0.478</td>
</tr>
<tr>
<td>LightCTS</td>
<td>0.465</td>
<td>0.477</td>
</tr>
<tr>
<td>Cross-LKTCN</td>
<td>0.455</td>
<td>0.461</td>
</tr>
<tr>
<td>PatchMixer</td>
<td>0.463</td>
<td>0.445</td>
</tr>
<tr>
<td>ModernTCN</td>
<td>0.471</td>
<td>0.457</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Li et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] proposed a deep learning-based prediction method, DLSF, which converts time series data into images for processing. For feature extraction, they designed a two-branch convolutional neural network. Finally, a linear autoregressive component is integrated to enhance robustness, making it suitable for dynamic cyclic or non-cyclic sequence prediction. For data prediction, they proposed a multilayer neural network to predict data changes. This method is used for power load prediction and takes into account external influences. The results showed good accuracy and efficiency on a load dataset from a major city in China. Borovykh et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] proposed a CNN model for financial predictive analytics inspired by the deep convolutional WaveNet architectural model. The model uses the ReLU activation function and employs parametric skip-connection conditioning to simplify and optimize the structure of the TSF forecasting model.</p>
<p>Although CNN performs well in prediction on some small-scale datasets, as the datasets become larger and more complex, CNN may appear to perform poorly on large datasets. To address this issue, Dong et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] proposed a model that combines CNN with the K-means clustering algorithm to achieve better accuracy and scalability. This method divides the dataset into different smaller sub-datasets using K-means and trains the CNN on the generated sub-datasets. The performance of the method on a large power dataset with more than 10,000 samples proves its effectiveness and high performance.</p>
<p>Current time series forecasting approaches predominantly focus on single-point prediction, failing to account for temporal interdependencies between forecasts at varying horizons, which consequently constrains their predictive performance. To bridge this gap, Cheng et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] introduced the Multi-Level Conformational Neural Network (MLCNN), an innovative multi-task learning architecture inspired by human predictive cognition. MLCNN uniquely integrates: (1) shared feature extraction across temporal scales, (2) dynamic fusion of multi-horizon predictive mechanisms, and (3) explicit modeling of cross-horizon interaction patterns&#x2014;collectively enhancing forecast accuracy through synergistic temporal relationship learning.</p>
<p>Generic models used to solve the TSF problem do not take into account the specificity between time series data well. To solve this problem, Liu et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] proposed Sample Convolution and Interaction Network (SCINet) to address the issue of generic models not accounting for the specificity between time series data. SCINet uses a downsampled convolutional interaction framework to simulate complex dynamic time series for better prediction. The SCI-Block module in this network structure extracts the input time series data and its features into two subsequences through downsampling, allowing the network to better learn the complex and rich features of the input time series.</p>
<p>To address the dual challenges of computational complexity in Transformer architectures and their limited capacity for local feature extraction, Wang et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] developed the Multi-scale Isometric Convolution Network (MICN). This innovative framework integrates parallel convolutional branches to simultaneously capture: (i) fine-grained local patterns through isometric kernels, and (ii) global temporal dependencies via hierarchical feature fusion. The multi-scale design explicitly decouples short-range and long-range modeling, achieving superior efficiency while maintaining modeling fidelity compared to conventional attention-based approaches.</p>
<p>To enhance the modeling capacity for complex temporal patterns, Wu et al. [<xref ref-type="bibr" rid="ref-24">24</xref>] introduced TimesNet, an innovative framework that transforms 1D time series into multiple 2D tensors through periodic decomposition. This novel representation: (i) encodes intra-cycle variations along tensor columns, (ii) captures inter-cycle dynamics through tensor rows, and (iii) enables efficient 2D convolution operations for joint temporal pattern learning. The architecture demonstrates state-of-the-art performance across five benchmark time series analysis tasks by effectively leveraging both microscopic periodic fluctuations and macroscopic trend evolution through its unique 2D transformation paradigm.</p>
<p>To mitigate the computational inefficiency of existing deep learning models for correlated time series forecasting without sacrificing accuracy, Lai et al. [<xref ref-type="bibr" rid="ref-25">25</xref>] developed LightCTS&#x2014;a lightweight framework employing streamlined temporal-spatial operator stacking. Unlike conventional architectures with alternating layers, LightCTS adopts: (i) parallelized operator modules for reduced computational overhead, (ii) optimized feature interaction mechanisms through simplified tensor operations, and (iii) adaptive weight sharing across prediction horizons.</p>
<p>Luo and Wang [<xref ref-type="bibr" rid="ref-26">26</xref>] proposed a convolution-based network architecture, CrossLKTCN, which addresses the problem that existing methods mainly focus on cross-temporal dependencies and do not adequately consider cross-variate dependencies. Gong et al. [<xref ref-type="bibr" rid="ref-27">27</xref>] proposed a novel CNN-based model, PatchMixer, to address the problem of temporal information loss due to the alignment-agnostic mechanism of transformer-based methods. PatchMixer retains temporal information by using a variational convolutional structure, relying solely on depth-separable convolution and using a single-scale architecture to simultaneously extract local features and global correlations of temporal information. Luo and Wang [<xref ref-type="bibr" rid="ref-28">28</xref>] brought the CNN convolutional model back into the spotlight by adapting the traditional Temporal Convolutional Network (TCN) and modifying it into a more suitable model for time series tasks, namely ModernTCN. ModernTCN adopts a large convolutional kernel, enhancing the receptive field. The method also leverages the ability of convolution to capture dependencies between variables, using three sets of convolutions to cleverly realize the decoupling modeling of three relationships: temporal, channel, and variable.</p>
<p>In summary, they may perform less effectively in modeling global patterns and complex time-series structures, such as non-smoothness, seasonality, or periodicity.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>RNN-Based Methods</title>
<p>Recurrent Neural Networks (RNNs), first proposed by Elman [<xref ref-type="bibr" rid="ref-29">29</xref>], play a crucial role in time series forecasting (TSF) tasks. RNNs are characterized by their recurrent connectivity, which allows them to store past information and utilize it in the current time step by introducing recurrent dependencies on the temporal dimensions. The structure of the RNN is shown in <xref ref-type="fig" rid="fig-17">Fig. 17</xref>.</p>
<fig id="fig-17">
<label>Figure 17</label>
<caption>
<title>The structure of the RNN used for the TSF task</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_68024-fig-17.tif"/>
</fig>
<p>Despite the suitability of RNNs for processing time series information, they have limitations in handling long sequences of temporal data and suffer from the vanishing gradient problem. To address these issues, Graves [<xref ref-type="bibr" rid="ref-30">30</xref>] introduced Long Short-Term Memory (LSTM). LSTM efficiently captures and conveys long-term dependencies in time series by incorporating memory units and gating mechanisms, thereby avoiding the gradient vanishing problem inherent in RNNs. Although LSTM solves the issues of RNNs, its complex structure demands greater computational resources and results in slower network training. To simplify the model, Cho et al. [<xref ref-type="bibr" rid="ref-31">31</xref>] proposed the Gated Recurrent Unit (GRU) as a streamlined version of LSTM. GRU contains only two gating mechanisms: the update gate and the reset gate. The update gate controls the weights of past memories and current inputs, while the reset gate determines the impact of past memories on current inputs. Some TSF methods based on RNN models are described below, as shown in <xref ref-type="table" rid="table-3">Table 3</xref>.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>RNN-based model</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>Year</th>
<th>Type</th>
<th>Baseline</th>
</tr>
</thead>
<tbody>
<tr>
<td>LSTNet</td>
<td>2018</td>
<td>Multivariate</td>
<td>AR, LSVR, TRMF, LRidge, GP, VARMLP, RNN-GRU</td>
</tr>
<tr>
<td>DA-RNN</td>
<td>2017</td>
<td>Univariate</td>
<td>ARIMA, Encoder-Decoder, NARX RNN, Attention RNN, Input-Attn-RNN</td>
</tr>
<tr>
<td>MTNet</td>
<td>2019</td>
<td>Multivariate</td>
<td>AR, GP, LRidge, LSVR, VARMLP, DA-RNN, RNN-GRU</td>
</tr>
<tr>
<td>MH-TAL</td>
<td>2020</td>
<td>Multivariate</td>
<td>Benchmark, POS-RNN, MQ-RNN</td>
</tr>
<tr>
<td>Jung et al.</td>
<td>2020</td>
<td>Multivariate</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>MTSMFF</td>
<td>2020</td>
<td>Multivariate</td>
<td>ARIMA, SVR-RBF, RNN, CNN, LSTM, GRU, VARMA, SVR-Linear,<break/>Seq2Seq, Seq2Seq-ATT, Seq2Seq-BI</td>
</tr>
<tr>
<td>STAM</td>
<td>2020</td>
<td>Multivariate</td>
<td>SVR-RBF, LSTM-Att, DA-RNN, Enc-Dec</td>
</tr>
<tr>
<td>CRU</td>
<td>2022</td>
<td>Irregular</td>
<td>RKN, GRU, Latent ODE, ODE-RNN, GRU-ODE-B</td>
</tr>
<tr>
<td>WITRAN</td>
<td>2023</td>
<td>Univariate</td>
<td>TimesNet, MICN, PatchTST, FiLm, DLinear, FEDformer,<break/>Pyraformer, Informer, Transformer, Autoformer</td>
</tr>
<tr>
<td>Fang et al.</td>
<td>2023</td>
<td>Multivariate</td>
<td>LSTM</td>
</tr>
<tr>
<td>SutraNets</td>
<td>2023</td>
<td>Univariate</td>
<td>&#x2013;</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>As shown in <xref ref-type="table" rid="table-4">Table 4</xref>, we evaluate the overall performance of the model on the ETTh1 dataset with a prediction length of 720.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Performance on the ETTh1 dataset. The prediction lengths are 720</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>MAE</th>
<th>MSE</th>
</tr>
</thead>
<tbody>
<tr>
<td>LSTNet</td>
<td>1.791</td>
<td>1.787</td>
</tr>
<tr>
<td>DA-RNN</td>
<td>1.499</td>
<td>1.511</td>
</tr>
<tr>
<td>MTNet</td>
<td>1.208</td>
<td>1.288</td>
</tr>
<tr>
<td>MH-TAL</td>
<td>0.968</td>
<td>0.897</td>
</tr>
<tr>
<td>Jung et al.</td>
<td>0.765</td>
<td>0.771</td>
</tr>
<tr>
<td>MTSMFF</td>
<td>0.855</td>
<td>0.834</td>
</tr>
<tr>
<td>STAM</td>
<td>0.690</td>
<td>0.679</td>
</tr>
<tr>
<td>CRU</td>
<td>0.643</td>
<td>0.622</td>
</tr>
<tr>
<td>WITRAN</td>
<td>0.491</td>
<td>0.451</td>
</tr>
<tr>
<td>Fang et al.</td>
<td>0.487</td>
<td>0.475</td>
</tr>
<tr>
<td>SutraNets</td>
<td>0.441</td>
<td>0.467</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Lai et al. [<xref ref-type="bibr" rid="ref-32">32</xref>] developed the Long- and Short-term Time Series Network (LSTNet), which synergistically combines convolutional layers for local pattern extraction with recurrent units for trend modeling, demonstrating superior performance on real-world datasets exhibiting complex periodic behaviors. Complementing this approach, Qin et al. [<xref ref-type="bibr" rid="ref-33">33</xref>] proposed the Dual-Stage Attention-Based RNN (DARNN), incorporating: (i) an input attention mechanism for dynamic feature selection based on historical encoder states, and (ii) a temporal attention module specifically designed to model extended dependencies.</p>
<p>Recent innovations in temporal modeling have introduced memory-enhanced architectures to overcome limitations in capturing complex temporal dependencies. Chang et al. [<xref ref-type="bibr" rid="ref-34">34</xref>] developed the Memory Time-series Network (MTNet), incorporating: (i) a large memory module for long-term pattern retention, (ii) triple independent encoders for multi-scale feature extraction, and (iii) an interpretable attention-based autoregressive component. Parallelly, Fan et al. [<xref ref-type="bibr" rid="ref-35">35</xref>] proposed a multi-view framework employing bidirectional LSTM decoders with temporal attention mechanisms, which dynamically integrate historical patterns and future contextual information to enhance prediction accuracy. These approaches collectively advance time series analysis through memory-augmented designs and attention-based temporal fusion, effectively addressing both long-term dependency capture and multi-view information integration challenges.</p>
<p>Jung et al. [<xref ref-type="bibr" rid="ref-36">36</xref>] proposed a predictive model for forecasting monthly PV generation using an RNN with an LSTM layer to process monthly time-series data. To address the dual challenges of multi-step and multivariate forecasting in classical models, Du et al. [<xref ref-type="bibr" rid="ref-37">37</xref>] developed MTSMFF, an encoder-decoder framework employing Bi-LSTM with attention mechanisms to simultaneously capture long-term temporal dependencies and cross-variable nonlinear interactions. Complementing this approach, Gangopadhyay et al. [<xref ref-type="bibr" rid="ref-38">38</xref>] introduced STAM, which integrates spatiotemporal attention with LSTM to explicitly model both temporal causality (through historical data constraints) and dynamic spatial correlations. These architectures collectively advance multivariate forecasting by: (i) leveraging bidirectional recurrent structures for enhanced temporal representation learning, and (ii) incorporating attention mechanisms to adaptively weight important features across both time and variable dimensions, as demonstrated through improved performance on complex real-world datasets with interdependent sensors and geographical distributed measurements.</p>
<p>Schirmer et al. [<xref ref-type="bibr" rid="ref-39">39</xref>] proposed Continuous Recurrent Units (CRUs) to handle irregular intervals between observations. CRUs assume a hidden state that evolves according to linear stochastic differential equations and are integrated into an encoder-decoder framework.</p>
<p>To address the difficulty of other methods in capturing different types of semantic information, Jia et al. [<xref ref-type="bibr" rid="ref-40">40</xref>] proposed a Waterwave Information Transfer (WIT) framework to capture long-term and short-term repetitive patterns through dual-grained information transfer. The framework uses a Horizontal Vertical Gated Selection Unit (HVGSU) to recursively fuse and select information, building global and local correlation models. The method improves prediction accuracy and handles challenges such as error accumulation and signal path distance through autoregressive generative modeling and cross-time, cross-subsequence modeling.</p>
<p>Overall, RNNs are effective in feature extraction and modeling temporal relationships in TSF tasks. However, using only RNN models may lead to gradient vanishing or explosion, so improved RNN models like LSTM and GRU are recommended for processing.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>GNN-Based Methods</title>
<p>Graph Neural Networks (GNNs) are a class of neural network models designed to handle graph-structured data and are widely used for spatio-temporal multivariate time series prediction. GNNs excel at capturing the dependencies of nodes and edges in the spatio-temporal dimension from time series data. By propagating information and aggregating features of neighboring nodes in the graph structure, GNNs can learn spatio-temporal representations of nodes and edges. GNNs can simultaneously process the information of multiple variables in the graph structure and perform multivariate time series prediction. By learning multidimensional feature representations at nodes and propagating and aggregating information in the graph structure, GNNs can synthesize the interactions and dependencies among different variables.</p>
<p>Several variants of GNNs have been developed to enhance their capabilities. For instance, the Graph Attention Network (GAT) [<xref ref-type="bibr" rid="ref-41">41</xref>] is. GAT uses attention weights to adaptively compute the importance of each node with its neighboring nodes, allowing it to more accurately aggregate the information of neighboring nodes. By learning the attention weights between different nodes, GAT can focus more on the neighboring nodes with importance.</p>
<p>Another common GNN variant is the Graph Convolutional Network (GCN) [<xref ref-type="bibr" rid="ref-42">42</xref>]. GCN updates the representation of a node using its neighbor information, similar to how convolutional neural networks perform convolutional operations on images. With multiple layers of graph convolution operations, GCNs can capture both local and global features of nodes, generating richer representations. GCNs are advantageous due to their simple structure and ease of implementation, and they have shown good performance in many graph data tasks.</p>
<p>Both GAT and GCN are important variants of GNNs, each with its own strengths. GAT is suitable for tasks that require accurate modeling of the importance between nodes, while GCN is suitable for tasks that require in-depth learning of local and global features of nodes.</p>
<p>Next, we introduce some GNN-based TSF methods, as shown in <xref ref-type="table" rid="table-5">Table 5</xref>.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>GNN-based model</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>Year</th>
<th>Type</th>
<th>Baseline</th>
</tr>
</thead>
<tbody>
<tr>
<td>STGCN</td>
<td>2018</td>
<td>SpatioTemporal</td>
<td>HA, LSVR, FC-LSTM, GCGRU, ARIMA, FNN,</td>
</tr>
<tr>
<td>ASTGCN</td>
<td>2019</td>
<td>SpatioTemporal</td>
<td>HA, VAR, LSTM, GRU, STGCN, GLU-STGCN, GeoMAN</td>
</tr>
<tr>
<td>SLC</td>
<td>2020</td>
<td>SpatioTemporal</td>
<td>HA, FNN, ARIMA, STGCN, DCRNN, GWN, SLCNN-P</td>
</tr>
<tr>
<td>MTGNN</td>
<td>2020</td>
<td>Multivariate</td>
<td>AR, GP, RNN-GRU, VARMLP, LSTNet-skip, TPA-LSTM</td>
</tr>
<tr>
<td>StemGNN</td>
<td>2020</td>
<td>Multivariate</td>
<td>FC-LSTM, SFM, N-BEATS, DCRNN, LSTNet, ST-GCN, TCN, DeepState, GraphWaveNet, DeepGLO</td>
</tr>
<tr>
<td>AutoSTG</td>
<td>2021</td>
<td>SpatioTemporal</td>
<td>HA, DCRNN, GBRT, GAT-Seq2Seq, ST-MetaNet</td>
</tr>
<tr>
<td>DMSTGCN</td>
<td>2021</td>
<td>SpatioTemporal</td>
<td>HA, VAR, LR, XGBoost, DCRNN, ASTGCN, GMAN, GWNet, MTGNN</td>
</tr>
<tr>
<td>D2STGNN</td>
<td>2022</td>
<td>SpatioTemporal</td>
<td>HA, VAR, DCRNN, STGCN, FC-LSTM, Grapg WaveNet, SVR, ASTGCN, MTGNN, STSGCN, DGCRN, GMAN</td>
</tr>
<tr>
<td>MAGNN</td>
<td>2022</td>
<td>Multivariate</td>
<td>Stock-LSTM, Stock-GAT, Event-NTN, News-ATT</td>
</tr>
<tr>
<td>FourierGNN</td>
<td>2023</td>
<td>Multivariate</td>
<td>VAR, SFM, LSTNet, TCN, DeppGLO, Reformer, Informer, Autoformer, FEDformer,Graph WaveNet, StemGNN, MTGNN, AGCRN</td>
</tr>
<tr>
<td>TPGNN</td>
<td>2023</td>
<td>SpatioTemporal</td>
<td>ARIMA, FC-LSTM, STGCN, DCRNN, StemGNN, Graph WaveNet, Informer, MTGNN</td>
</tr>
<tr>
<td>MSGNet</td>
<td>2023</td>
<td>Multivariate</td>
<td>TimesNet, DLinear, Nlinear, MTGNN, Autoformer, Informer</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>As shown in <xref ref-type="table" rid="table-6">Table 6</xref>, we evaluate the overall performance of the model on the ETTh1 dataset with a prediction length of 720.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Performance on the ETTh1 dataset. The prediction lengths are 720</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>MAE</th>
<th>MSE</th>
</tr>
</thead>
<tbody>
<tr>
<td>STGCN</td>
<td>1.998</td>
<td>1.778</td>
</tr>
<tr>
<td>ASTGCN</td>
<td>1.199</td>
<td>1.191</td>
</tr>
<tr>
<td>SLC</td>
<td>1.208</td>
<td>1.088</td>
</tr>
<tr>
<td>MTGNN</td>
<td>1.120</td>
<td>1.185</td>
</tr>
<tr>
<td>StemGNN</td>
<td>0.965</td>
<td>0.911</td>
</tr>
<tr>
<td>AutoSTG</td>
<td>0.955</td>
<td>0.934</td>
</tr>
<tr>
<td>DMSTGCN</td>
<td>0.890</td>
<td>0.875</td>
</tr>
<tr>
<td>D2STGNN</td>
<td>0.843</td>
<td>0.882</td>
</tr>
<tr>
<td>MAGNN</td>
<td>0.791</td>
<td>0.777</td>
</tr>
<tr>
<td>FourierGNN</td>
<td>0.681</td>
<td>0.679</td>
</tr>
<tr>
<td>TPGNN</td>
<td>0.541</td>
<td>0.515</td>
</tr>
<tr>
<td>MSGNet</td>
<td>0.488</td>
<td>0.494</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Yu et al. [<xref ref-type="bibr" rid="ref-43">43</xref>] proposed a spatio-temporal graph convolutional network (STGCN) to address the limitations of traditional methods in medium- and long-term traffic prediction. STGCN models the problem with a complete convolutional structure on the graph, efficiently capturing comprehensive spatio-temporal correlations and outperforming other models on various traffic datasets by modeling the multiscale traffic network. Guo et al. [<xref ref-type="bibr" rid="ref-44">44</xref>] proposed an attention-based spatio-temporal graph convolution network (ASTGCN). ASTGCN consists of three independent components to model the three temporal attributes of the traffic flow: near-term dependency, daily cycle dependency, and weekly cycle dependency [<xref ref-type="bibr" rid="ref-45">45</xref>]. Each component contains a spatio-temporal attention mechanism for capturing dynamic spatio-temporal correlations and spatio-temporal convolution in traffic data, while graph convolution is employed to capture spatial patterns and ordinary standard convolution to describe temporal features [<xref ref-type="bibr" rid="ref-45">45</xref>]. Zhang et al. [<xref ref-type="bibr" rid="ref-46">46</xref>] identified three challenges in traffic data: (1) traffic data is physically associated with a road network and should be formatted as a traffic graph rather than a plain grid-like tensor, (2) traffic data has strong spatial dependencies, and (3) traffic data has strong time dependencies. To address these issues, they proposed a network framework called Structure Learning Convolution (SLC) [<xref ref-type="bibr" rid="ref-47">47</xref>].</p>
<p>Wu et al. [<xref ref-type="bibr" rid="ref-48">48</xref>] proposed a generalized graph neural network framework for multivariate time series data, allowing the model to automatically extract unidirectional relationships between variables and incorporate external factors such as variable attributes [<xref ref-type="bibr" rid="ref-49">49</xref>]. Cao et al. [<xref ref-type="bibr" rid="ref-50">50</xref>] proposed a Spectral Temporal Graph Neural Network (StemGNN) that captures both intra-sequence temporal correlation and inter-sequence correlation to improve the accuracy of multivariate time series prediction.</p>
<p>Pan et al. [<xref ref-type="bibr" rid="ref-51">51</xref>] proposed AutoSTG, a new network for automatic spatio-temporal graph forecasting. With the aim of exploring the inherent dynamics in traffic data such as traffic speed, traffic volume, and multifaceted spatio-temporal features for better prediction of traffic speed, Han et al. [<xref ref-type="bibr" rid="ref-52">52</xref>] proposed a dynamic graph construction method based on Dynamic Graph Neural Networks (DGNN) to learn the time-specific spatial correlations of road segments. Shao et al. [<xref ref-type="bibr" rid="ref-53">53</xref>] proposed a decoupled spatio-temporal framework (DSTF) to address the problem of traffic data containing both diffuse and intrinsic signals. DSTF separates these signals in a data-driven manner and processes them separately. They also proposed Decoupled Dynamic Spatio-Temporal Graph Neural Network (D2STGNN), which captures spatio-temporal correlations and has a dynamic graph learning module [<xref ref-type="bibr" rid="ref-53">53</xref>].</p>
<p>Financial time series analyses are usually characterized by multi-modal flows and overshooting lag effects, and the financial industry needs predictive models that are interpretable and compatible. Based on these needs, Cheng et al. [<xref ref-type="bibr" rid="ref-54">54</xref>] proposed a multimodal graph neural network (MAGNN) for financial time series forecasting. MAGNN constructs a heterogeneous graph network with sources in the financial knowledge graph as nodes and relations as edges [<xref ref-type="bibr" rid="ref-55">55</xref>].</p>
<p>Latent variable correlations for multivariate time series forecasting are more complex, and the dominant approach using GNNs is to represent the correlations as static graphs, but this approach can lead to significant bias due to the fact that correlations in multivariate time data are constantly changing over time. To address this problem, Liu et al. [<xref ref-type="bibr" rid="ref-56">56</xref>] proposed a temporal polynomial graph neural network (TPGNN) for accurate multivariate time series prediction. TPGNN starts with a static matrix to capture overall correlation and constructs a matrix polynomial for each time step using time-varying coefficients and a matrix basis. Cai et al. [<xref ref-type="bibr" rid="ref-57">57</xref>] proposed MSGNet, it employs a self-attention mechanism and an adaptive hybrid graph convolutional layer to learn different inter-sequence correlations within each time scale.</p>
<p>Overall, GNNs enhance the accuracy and robustness of time series forecasts by capturing the spatio-temporal relationships and correlations between multivariate variables using the characteristic representations of nodes and edges in the graph structure.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Transformer-Based Methods</title>
<p>The Transformer architecture [<xref ref-type="bibr" rid="ref-58">58</xref>], developed for natural language processing, has become a pivotal framework for temporal modeling due to its self-attention mechanism (<xref ref-type="fig" rid="fig-18">Fig. 18</xref>). Its encoder-decoder structure operates through: (i) hierarchical feature abstraction in the encoder via multi-head attention, and (ii) autoregressive sequence generation in the decoder. This design fundamentally addresses the long-term dependency learning limitations of RNNs by eliminating recurrent connections, thereby preventing gradient vanishing/explosion issues while enabling parallel processing of entire sequences. The model&#x2019;s capability to simultaneously attend to all temporal positions through attention weights allows direct capture of both local patterns and global trends in time series data, making it particularly effective for applications requiring modeling of extended temporal contexts.</p>
<fig id="fig-18">
<label>Figure 18</label>
<caption>
<title>The structure of the Transformer used for the TSF task</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_68024-fig-18.tif"/>
</fig>
<p>Moreover, compared to RNN models, the self-attention mechanism in Transformer allows for better parallelism and higher computational efficiency. Some applications of Transformer in TSF tasks are presented next, as shown in <xref ref-type="table" rid="table-7">Table 7</xref>.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Transformer-based model</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>Year</th>
<th>Type</th>
<th>Baseline</th>
</tr>
</thead>
<tbody>
<tr>
<td>LogTrans</td>
<td>2020</td>
<td>Univariate</td>
<td>ARIMA, ETS, TRMF, DeepAR, DeepState</td>
</tr>
<tr>
<td>trafficBERT</td>
<td>2021</td>
<td>Multivariate</td>
<td>ARIMA, SAE, LSTM, FC-LSTM, FC-GRU</td>
</tr>
<tr>
<td>AST</td>
<td>2020</td>
<td>Univariate</td>
<td>ARIMA, ETS, TRMF, DeepAR, DSSM, ConvTrans</td>
</tr>
<tr>
<td>SpringNet</td>
<td>2020</td>
<td>Multivariate</td>
<td>DeepAR, Transformer</td>
</tr>
<tr>
<td>Informer</td>
<td>2021</td>
<td>Multivariate</td>
<td>LogTrans, Reformer, LSTM, DeepAR, ARIMA, Prophet</td>
</tr>
<tr>
<td>TFT</td>
<td>2021</td>
<td>Multivariate</td>
<td>ARIMA, ETS, TRMF, DeepAR, DSSM, ConvTrans, Seq2Seq,<break/>MQRNN</td>
</tr>
<tr>
<td>SSDNet</td>
<td>2021</td>
<td>Multivariate</td>
<td>Persistence, SARIMAX, Prophet, DeepAR, DeepSSM,<break/>N-BEATS, LogSparse Transformer, Informer</td>
</tr>
<tr>
<td>Autoformer</td>
<td>2021</td>
<td>Multivariate</td>
<td>Informer, LogTrans, Reformer, LSTNet, LSTM, TCN</td>
</tr>
<tr>
<td>Aliformer</td>
<td>2021</td>
<td>Univariate</td>
<td>Informer, LogTrans, LSTM, LSTNet</td>
</tr>
<tr>
<td>Pyraformer</td>
<td>2022</td>
<td>Multivariate</td>
<td>Informer, LogTrans, Longformer, Reformer, etc</td>
</tr>
<tr>
<td>FEDformer</td>
<td>2022</td>
<td>Multivariate</td>
<td>Autoformer, Informer, LogTrans, Reformer</td>
</tr>
<tr>
<td>Triformer</td>
<td>2022</td>
<td>Multivariate</td>
<td>Reformer, LogTrans, StemGNN, AGCRN, Informer, Autoformer</td>
</tr>
<tr>
<td>Quatformer</td>
<td>2022</td>
<td>Multivariate</td>
<td>Autoformer, Informer, LogTrans, Reformer, LSTM, TCN</td>
</tr>
<tr>
<td>Crossformer</td>
<td>2023</td>
<td>Multivariate</td>
<td>LSTM, LSTNet, MTGNN, Transformer, Informer, Autoformer,<break/>Pyraformer, FEDformer</td>
</tr>
<tr>
<td>Airformer</td>
<td>2023</td>
<td>Spatio<break/>Temporal</td>
<td>HA, VAR, DCRNN, STGCN, GWNET, MTGNN, ASTGCN,<break/>GMAN, STTN, DeepAir, <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mi>P</mml:mi><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mn>2.5</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>-GNN, GAGNN</td>
</tr>
<tr>
<td>PatchTST</td>
<td>2023</td>
<td>Univariate</td>
<td>DLinear, FEDformer, Autoformer, Informer, Pyraformer, LogTrans</td>
</tr>
<tr>
<td>Detformer</td>
<td>2023</td>
<td>Multivariate</td>
<td>Informer, EDLSTM, EDGruAtt</td>
</tr>
<tr>
<td>Scaleformer</td>
<td>2023</td>
<td>Multivariate</td>
<td>FEDformer, Reformer, Performer, Informer, Autoformer</td>
</tr>
<tr>
<td>Conformer</td>
<td>2023</td>
<td>Multivariate</td>
<td>GRU, LSTNet, N-Beats, Reformer, Longformer,<break/>LogTrans, Informer, Autoformer, TS2Vec</td>
</tr>
<tr>
<td>PDformer</td>
<td>2023</td>
<td>SpatioTemporal</td>
<td>STResNet, DMVSTNet, DSAN, VAR, SVR, DCRNN, STGCN,<break/>GWNET, MTGNN, STSGCN, STFGNN, STGODE, STGNCDE,<break/>STTN, GMAN, TFormer, ASTGNN</td>
</tr>
<tr>
<td>GCformer</td>
<td>2023</td>
<td>Multivariate</td>
<td>PatchTST, MICN, FEDformer, Autoformer, S4, Informer,<break/>LogTrans, DLinear</td>
</tr>
<tr>
<td>Sageformer</td>
<td>2023</td>
<td>Multivariate</td>
<td>Crossformer, MTGNN, LSTNet, Transformer, Informer,<break/>Autoformer, Non-stationary Transformer</td>
</tr>
<tr>
<td>Difformer</td>
<td>2023</td>
<td>Multivariate</td>
<td>Autoformer, Informer, transformer, DeepAR,<break/>ARIMA, UniTS</td>
</tr>
<tr>
<td>DSformer</td>
<td>2023</td>
<td>Multivariate</td>
<td>PatchTST, Crossformer, TimesNet, DLinear, FEDformer,<break/>Pyraformer, Autoformer, Informer</td>
</tr>
<tr>
<td>iTransformer</td>
<td>2023</td>
<td>Multivariate</td>
<td>Autoformer, FEDformer, Stationary, Crossformer,<break/>PatchTST, DLinear, TiDE, RLinear, SCINet, TimesNet</td>
</tr>
<tr>
<td>MASTER</td>
<td>2023</td>
<td>Multivariate</td>
<td>XGBoost, LSTM, GRU, TCN, Transformer, GAT, DTML</td>
</tr>
<tr>
<td>CARD</td>
<td>2023</td>
<td>Multivariate</td>
<td>FEDformer, ETSformer, FilM, LightTS, MICN,<break/>TimesNet, Dlinear, PatchSTS</td>
</tr>
<tr>
<td>Contiformer</td>
<td>2023</td>
<td>Irregular</td>
<td>GRU, ODE-RNN,CADN, Neural CDE, S5, TST, mTAN</td>
</tr>
<tr>
<td>Basisformer</td>
<td>2024</td>
<td>Multivariate</td>
<td>FEDformer, Autoformer, Pyraformer, DLinear,<break/>TCN, N-Hits, FiLM</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Li et al. [<xref ref-type="bibr" rid="ref-59">59</xref>] found that Transformer has the problems of limiting diagnosis and memory bottleneck, and to solve these two problems, The authors introduced a novel Convolutional Self-Attention mechanism that revolutionizes traditional attention by employing causal convolutions to generate queries and keys, thereby effectively integrating localized contextual information into the attention computation process. Then, they proposed the LogSparse Transformer with a memory cost of only <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>L</mml:mi><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, which improves the prediction accuracy of time series with fine-grained and strong long-term dependencies under a limited memory budget [<xref ref-type="bibr" rid="ref-60">60</xref>].</p>
<p>Jin et al. [<xref ref-type="bibr" rid="ref-61">61</xref>] proposed trafficBERT based on the BERT [<xref ref-type="bibr" rid="ref-13">13</xref>] model. This model captures time series information by using multi-head self-attention instead of the commonly used RNN [<xref ref-type="bibr" rid="ref-61">61</xref>]. The model requires only information about traffic speeds and roads on days of the week for prediction and does not require information about the flow on neighboring roads at the current moment, which has few application limitations.</p>
<p>Tansformer is deficient in facing long sequence time prediction. In terms of this problem, Zhou et al. [<xref ref-type="bibr" rid="ref-62">62</xref>] proposed the Informer model based on the Transformer encoder-decoder structure. The Informer model can give all the required long sequence prediction results at one time, instead of adopting the method of multiple prediction for prediction. Informer first proposes the ProbSparse self-attention mechanism, which can effectively handle longer sequence input data by replacing the traditional Transformer self-attention in the encoder part by employing the multi-head sparse self-attention [<xref ref-type="bibr" rid="ref-63">63</xref>]. Secondly, it proposed a self-attention refining mechanism, which can greatly reduce the number of layers of the network and improve the robustness of the layer stacking part by extracting the self-attention distillation part of the dominant attention. In addition, the decoder part of the model sets the predicted sequence and the subsequent data to 0 for data masking, and the sequence input requires only one forward step, which effectively avoids error accumulation. Informer introduces sparse bias in the self-attention model, as well as Logsparse masking, which reduces the computational complexity of the traditional Transformer model from <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>L</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> to <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>L</mml:mi><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>L</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
<p>Lim et al. [<xref ref-type="bibr" rid="ref-64">64</xref>] proposed the Time Fusion Transformer (TFT), which uses a recurrent layer for localization and an interpretable self-attentive layer to capture long-term dependencies of the input sequence data. The model can be used for multilevel prediction with interpretability and high performance. Lin et al. [<xref ref-type="bibr" rid="ref-65">65</xref>] proposed the State Space Decomposition Neural Network (SSDNet) that combines Transformer and State Space Modeling (SSM) with both the high performance of deep learning and the interpretability of SSM [<xref ref-type="bibr" rid="ref-65">65</xref>]. It uses the Transformer architecture to learn temporal patterns and directly estimates the parameters of the SSM [<xref ref-type="bibr" rid="ref-65">65</xref>]. Wu et al. [<xref ref-type="bibr" rid="ref-66">66</xref>] proposed Autoformer based on the Transformer model. which transforms Transformer into a decomposition prediction architecture by embedding decomposition blocks as internal operators into the network structure. The model replaces Transformer&#x2019;s self-attention mechanism with autocorrelation mechanism, which can discover sequence similarity from sequence periodicity. Qi et al. [<xref ref-type="bibr" rid="ref-67">67</xref>] proposed a bi-directional Transformer, Aliformer, for dealing with time series sales forecasting problems in e-commerce. The model is designed with a knowledge-guided self-attentive layer and utilizes historical information, current factors and future knowledge to predict future data changes.</p>
<p>Liu et al. [<xref ref-type="bibr" rid="ref-68">68</xref>] proposed Pyraformer to capture a wide range of temporal dependencies [<xref ref-type="bibr" rid="ref-69">69</xref>]. The method introduces a Pyramid Attention Module (PAM) [<xref ref-type="bibr" rid="ref-70">70</xref>] in which a cross-scale tree structure generalizes features at different resolutions, while intra-scale neighbor connections model different ranges of temporal dependencies [<xref ref-type="bibr" rid="ref-69">69</xref>]. Recent advances in Transformer-based time series forecasting have addressed two critical challenges: computational efficiency and global pattern capture. Pyraformer demonstrates superior performance in empirical evaluations, achieving optimal prediction accuracy with minimal computational overhead&#x2014;particularly for long-sequence scenarios in both single-step and extended-horizon forecasting tasks. Building on this progress, Zhou et al. [<xref ref-type="bibr" rid="ref-71">71</xref>] introduced FEDformer, which enhances the standard Transformer architecture through: (i) a frequency-domain decomposition strategy to reduce computational complexity, and (ii) specialized Fourier and wavelet enhancement modules that replace conventional attention mechanisms. These innovations collectively enable more efficient modeling of global temporal structures while maintaining competitive predictive performance, as validated through comprehensive benchmarks on large-scale datasets.</p>
<p>A new attention-based Transformer model, Triformer, was proposed by Cirstea et al. [<xref ref-type="bibr" rid="ref-72">72</xref>]. They first proposed an attention mechanism called Patch Attention and designed a new triangular structure that stacks the attention layers, resulting in a significant reduction in the number of layers. In addition, they propose a lightweight method for modeling specific variables by introducing different projection matrices, which can capture different temporal patterns and improve prediction accuracy.</p>
<p>Chen et al. [<xref ref-type="bibr" rid="ref-73">73</xref>] proposed an innovative time series forecasting framework, Quatformer, which handles complex periodic patterns by introducing quaternion-based Learned Rotational Attention (LRA), and tackles the challenges of long-term dependencies and dot-product attention through trend normalization and global memory decoupling. In evaluations on multiple real-world datasets, Quatformer demonstrates its strengths in time series forecasting by improving performance by an average of 8.1% compared to state-of-the-art benchmark models, with up to 18.5% MSE improvement. Recent advancements in Transformer-based time series forecasting have introduced innovative architectures to address spatiotemporal dependencies. Crossformer effectively preserves temporal and dimensional information through its two-stage attention (TSA) layer, which captures cross-time and cross-variable relationships, while its hierarchical encoder-decoder (HED) leverages multi-scale representations for enhanced prediction accuracy, outperforming existing methods across six real-world benchmarks. Similarly, Liang et al. [<xref ref-type="bibr" rid="ref-74">74</xref>] developed Airformer for large-scale air quality forecasting, employing a two-stage framework that combines deterministic spatiotemporal attention with stochastic uncertainty modeling to achieve 5%&#x2013;8% error reduction in 72-h predictions across thousands of Chinese monitoring stations. These models demonstrate the growing capability of attention-based architectures to handle complex real-world forecasting tasks through specialized mechanisms for dependency modeling and multi-scale feature utilization.</p>
<p>Nie et al. [<xref ref-type="bibr" rid="ref-75">75</xref>] present PatchTST, it significantly improves the accuracy of long-term forecasts by splitting time series into patches with shared embeddings and weights. Meng et al. [<xref ref-type="bibr" rid="ref-76">76</xref>] proposed Detformer, which solves the problem that the Transformer-based method does not have the ability of temporal modeling resulting in the model not being directly applied. The method proposes a dual-feedback sparse attention mechanism to improve the stability of heuristic sparse attention. Also, they designed a time-dependent extraction mechanism to model the perspective of the attention index [<xref ref-type="bibr" rid="ref-76">76</xref>]. In addition, they proposed an algorithm to eliminate data noise so as to optimize the spatio-temporal modeling. Shabani et al. [<xref ref-type="bibr" rid="ref-77">77</xref>] propose ScaleFormer, a generalized multi-scale Transformer framework that enhances existing architectures through three key innovations: (1) iterative multi-scale refinement of predictions using weight-sharing mechanisms to maintain parameter efficiency, (2) strategic architectural modifications for improved temporal representation learning, and (3) a novel normalization scheme specifically optimized for multi-scale processing. This unified approach demonstrates consistent performance improvements across diverse Transformer variants and datasets while introducing minimal computational overhead, effectively addressing the trade-off between model capacity and efficiency in time series forecasting tasks. The framework&#x2019;s adaptability is further evidenced by its ability to enhance both local pattern capture and global trend modeling through its hierarchical refinement process.</p>
<p>Li et al. [<xref ref-type="bibr" rid="ref-78">78</xref>] proposed Conformer based on Transformer, which is specialized for long-term time series forecasting applications (LTTF) such as wind supply planning. The method achieves higher information utilization and accuracy and is capable of generating reliable forecasts with uncertainty quantification by introducing innovative designs such as an encoder-decoder architecture, a regularized flow module, and explicitly modeling the correlation and dynamics of the time series. Jiang et al. [<xref ref-type="bibr" rid="ref-79">79</xref>] proposed PDformer for accurate traffic flow prediction. Compared with traditional GNN models, PDformer features breakthrough innovations in modeling the spatio-temporal dependencies of urban traffic data, including dynamic spatial dependency capture, long-range spatial dependency modeling, and consideration of propagation time delay of traffic conditions. After extensive experimental validation, PDformer is not only state-of-the-art in terms of performance, but also competitively computationally efficient and makes its model highly interpretable by visualizing spatio-temporal attention maps.</p>
<p>Zhao et al. [<xref ref-type="bibr" rid="ref-80">80</xref>] proposed GCformer, a model that combines a global convolutional branch and a local Transformer branch, to address the limitations of Transformer in the prediction of long-input time series. Zhang et al. [<xref ref-type="bibr" rid="ref-81">81</xref>] proposed Sageformer. As a graph-enhanced Transformer model, Sageformer efficiently captures complex relationships within and between sequences and reduces redundant information.</p>
<p>Li et al. [<xref ref-type="bibr" rid="ref-82">82</xref>] present an effective and efficient Transformer architecture called DifFormer for performing various time series analysis tasks. Compared to previous Transformer variants, DifFormer employs a novel multi-resolution differencing mechanism that is capable of progressively and adaptively highlighting subtle but meaningful variations and is flexible enough to capture periodic or cyclic patterns [<xref ref-type="bibr" rid="ref-82">82</xref>]. Yu et al. [<xref ref-type="bibr" rid="ref-83">83</xref>] proposed a two-sampling transformer model called DSformer for long-term forecasting of multivariate time series.</p>
<p>Liu et al. [<xref ref-type="bibr" rid="ref-84">84</xref>] proposed iTransformer, a time series forecasting model based on the Transformer architecture. By applying attention and feedforward networks on the inverted dimension, iTransformer is able to better capture correlations between multivariate variables and learn nonlinear representations. Experimental results show that iTransformer achieves state-of-the-art performance on real datasets, providing better performance and generalization capabilities for Transformer models in time series forecasting. Li et al. [<xref ref-type="bibr" rid="ref-85">85</xref>] proposed MASTER (MArkert-Guided Stock TransformER) for stock price forecasting, aiming to solve the forecasting challenges caused by high volatility in the stock market. Unlike existing methods, MASTER efficiently models complex stock correlations by considering both instantaneous and intertemporal stock correlations and using market information for automatic feature selection.</p>
<p>Wang et al. [<xref ref-type="bibr" rid="ref-86">86</xref>] proposed the Channel Aligned Robust Blend Transformer (CARD) to address limitations of channel-independent Transformers in time series forecasting. The model introduces three key innovations: (1) a channel-aligned attention mechanism that simultaneously captures inter-variable dependencies and temporal correlations, (2) a multi-scale token mixing module that generates hierarchical representations at varying resolutions, and (3) a novel robust loss function incorporating uncertainty-weighted temporal importance to prevent overfitting. The framework effectively balances the modeling of cross-channel relationships with temporal dynamics while maintaining robustness through its specialized loss formulation.</p>
<p>Recent advances in Transformer-based time series forecasting have introduced innovative architectures to address key challenges in irregular and continuous-time data modeling. Chen et al. [<xref ref-type="bibr" rid="ref-87">87</xref>] developed ContiFormer, which integrates Neural ODE&#x2019;s continuous dynamics modeling with Transformer attention mechanisms to effectively handle irregular temporal patterns, demonstrating superior performance in continuous-time scenarios. Complementing this approach, Ni et al. [<xref ref-type="bibr" rid="ref-88">88</xref>] proposed BasisFormer, an interpretable framework that leverages self-learned basis functions through adaptive self-supervised learning. By employing bidirectional cross-attention to compute similarity coefficients between historical patterns and basis functions, BasisFormer achieves state-of-the-art performance with 11.04% to 15.78% improvements in univariate and multivariate forecasting tasks respectively across six benchmark datasets. These models collectively advance time series forecasting by combining the relational modeling strengths of Transformers with specialized mechanisms for continuous-time dynamics (ContiFormer) and interpretable pattern decomposition (BasisFormer), while addressing both irregular sampling and prediction accuracy challenges.</p>
<p>In summary, Transformer can capture long-term dependencies well and handle multivariate time series data, and the TSF method based on it has good robustness and generalization ability. It is worth noting that the processing effect of Transformer may be different when facing different time series datasets. Therefore, when dealing with the TSF task, the model selection should be based on the characteristics of the data.</p>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Recent Advances</title>
<p><bold>Diffusion-Based Forecasting for Time Series:</bold> Diffusion models learn data distributions by gradually adding and removing noise, and have recently been successfully applied to time series forecasting. These methods are especially suitable for high-uncertainty domains such as finance and healthcare, as they can generate multiple possible future trajectories, thereby quantifying prediction uncertainty. For example, TimeGrad uses RNNs to encode time series features and then generates probabilistic forecasts through the diffusion process. Recent research like DiffTS further combines diffusion models with Transformers, leveraging self-attention to capture long-term dependencies while retaining the generative capabilities of diffusion models. Additionally, it highlights the need for interpretability in finance, so the generative process of diffusion models should be combined with attention mechanisms or other interpretability tools to enhance trustworthiness.</p>
<p><bold>Hybrid Classical Statistical and Deep Learning Pipelines:</bold> Hybrid approaches combine traditional statistical models (such as ARIMA, GARCH) with deep learning models (such as LSTM, Transformer), improving forecasting performance while maintaining interpretability. For example, DeepAR uses autoregressive models to handle linear trends and then applies RNNs to learn nonlinear residuals; N-BEATS dynamically adjusts model weights via interpretable basis expansion modules. Such methods naturally satisfy the explainable AI (XAI) requirements described, since the statistical components provide clear parameter explanations and the deep learning parts can be further analyzed with tools like SHAP.</p>
<p><bold>Quantile Regression Combined with Transformers:</bold> Quantile regression directly predicts intervals at different confidence levels (e.g., 5%, 50%, 95% quantiles), providing richer information for financial risk management and decision-making. Transformers, with their powerful sequence modeling capabilities, serve as an ideal framework. For instance, Informer employs quantile attention heads to output multiple quantile forecasts simultaneously, avoiding strong assumptions about data distributions inherent in traditional methods. FEDformer further decomposes time series in the frequency domain and performs quantile regression on each subcomponent, better handling periodicity and sudden events.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Datasets and Performance Evaluation Metrics</title>
<p>This section provides an overview of datasets and performance evaluation metrics commonly used in time series forecasting (TSF) tasks. <xref ref-type="sec" rid="s4_1">Section 4.1</xref> introduces benchmark datasets across various application domains, summarized in tabular form.</p>
<p><xref ref-type="sec" rid="s4_2">Section 4.2</xref> discusses strategies for training and validating TSF models, with a focus on splitting data into training, validation, and test sets, as well as cross-validation techniques tailored for time series. <xref ref-type="sec" rid="s4_3">Section 4.3</xref> explains the importance of using specialized cross-validation techniques like TimeSeriesSplit and Walk-Forward Validation for accurately assessing time series models while maintaining the integrity of temporal dependencies. <xref ref-type="sec" rid="s4_4">Section 4.4</xref> addresses strategies for handling data imbalances and outliers in time series, emphasizing preprocessing techniques like robust scaling, outlier detection, and careful treatment of rare events to ensure accurate model performance. <xref ref-type="sec" rid="s4_5">Section 4.5</xref> highlights the role of data augmentation in time series forecasting, discussing methods like time warping, jittering, and bootstrapping to artificially expand the dataset and improve model generalization. (R2.18: <xref ref-type="sec" rid="s4_2">Sections 4.2</xref>&#x2013;<xref ref-type="sec" rid="s4_5">4.5</xref> are additions made in response to the reviewer&#x2019;s comments). <xref ref-type="sec" rid="s4_6">Section 4.6</xref> discusses widely adopted performance evaluation metrics for assessing TSF models.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Datasets for TSF</title>
<p>The selection and preparation of datasets are critical for algorithm validation, model comparison, and research analysis in TSF. Prior to utilization, datasets typically undergo preprocessing steps such as subset selection, noise reduction, missing value imputation. When addressing real-world problems, it is essential to select appropriate prediction models and algorithms based on the specific characteristics and requirements of the dataset. Blindly adopting state-of-the-art algorithms without considering the problem context may lead to suboptimal results. Researchers should carefully evaluate the number of feature variables, the required prediction horizon, and the scale of the dataset (e.g., the order of magnitude of records) when designing TSF solutions. Below, we describe datasets commonly used in TSF tasks across different application domains.
<list list-type="simple">
<list-item><label>1.</label><p><bold>Industrial Energy Datasets</bold></p>
<p>In the industrial energy sector, TSF plays a pivotal role in long-term strategic resource planning. It enables the prediction of future energy demand (e.g., electricity, oil, and natural gas), facilitating optimized production and supply planning. Additionally, TSF assists power companies in forecasting future power generation to ensure stable and adequate supply. These capabilities have broad applications, helping organizations and governments improve planning, mitigate risks, enhance efficiency, and achieve sustainable development goals. <xref ref-type="table" rid="table-8">Table 8</xref> summarizes key datasets relevant to industrial energy forecasting.</p></list-item>

<list-item><label>2.</label><p><bold>Financial Datasets</bold></p>
<p>TSF is extensively applied in finance, including the prediction of economic cycles, fiscal trends, and stock market behavior. These forecasts provide valuable decision support for financial traders, businesses, and policymakers. In stock markets, TSF models predict price trends and fluctuations, aiding investors in developing robust investment strategies. Beyond market analysis, TSF supports financial institutions in revenue and expenditure planning, loan risk assessment, and interest rate forecasting, thereby informing monetary policy formulation. <xref ref-type="table" rid="table-9">Table 9</xref> presents a summary of widely used financial datasets.</p>
</list-item>
<list-item><label>3.</label><p><bold>Meteorological Datasets</bold></p>
<p>In meteorology, TSF is employed for long-term climate trend prediction, natural disaster early warning, and marine weather forecasting. These applications provide critical decision support for agriculture, marine transportation, and disaster management, while also contributing to national climate change adaptation strategies. <xref ref-type="table" rid="table-10">Table 10</xref> summarizes key meteorological datasets used in TSF research.</p>
</list-item>
</list><table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Industrial energy dataset</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col/>
<col/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Ref.</th>
<th>Time range</th>
<th align="center">Time interval</th>
<th>Information</th>
</tr>
</thead>
<tbody>
<tr>
<td>ETT</td>
<td>[<xref ref-type="bibr" rid="ref-62">62</xref>,<xref ref-type="bibr" rid="ref-66">66</xref>]</td>
<td>2016.7&#x2013;2018.7</td>
<td>Hour, 15 min</td>
<td>ETT dataset records the load and oil temperature of power transformers.</td>
</tr>
<tr>
<td>Electricity</td>
<td>[<xref ref-type="bibr" rid="ref-66">66</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>,<xref ref-type="bibr" rid="ref-89">89</xref>]</td>
<td>2011&#x2013;2014</td>
<td>15 min</td>
<td>Electricity dataset records the electricity consumption of 321 customers.</td>
</tr>
<tr>
<td>Power<break/>consumption</td>
<td>[<xref ref-type="bibr" rid="ref-90">90</xref>]</td>
<td>2006.12&#x2013;2010.11</td>
<td>Minute</td>
<td>This dataset records the electricity consumption of a household over a period of nearly 4 years.</td>
</tr>
<tr>
<td>Wind</td>
<td>[<xref ref-type="bibr" rid="ref-91">91</xref>]</td>
<td>1986&#x2013;2015</td>
<td>Hour</td>
<td>This dataset records hourly<break/>estimates of energy potential as a percentage of the maximum output of power plants for a European region for the period 1986&#x2013;2015.</td>
</tr>
<tr>
<td>Hanergy</td>
<td>[<xref ref-type="bibr" rid="ref-92">92</xref>]</td>
<td>2011.1.1&#x2013;2016.12.31</td>
<td>Day</td>
<td>This dataset records solar power generation data from two photovoltaic plants in Alice Springs, Northern Territory, Australia.</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>Finance dataset</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col/>
<col/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Ref.</th>
<th>Time range</th>
<th align="center">Time interval</th>
<th>Information</th>
</tr>
</thead>
<tbody>
<tr>
<td>S&#x0026;P 500</td>
<td>[<xref ref-type="bibr" rid="ref-93">93</xref>],</td>
<td>2001.1&#x2013;2017.5</td>
<td>Day</td>
<td>This dataset records the daily<break/>S&#x0026;P 500 index from 2001.01&#x2013;2017.05.</td>
</tr>
<tr>
<td>Exchange rate</td>
<td>[<xref ref-type="bibr" rid="ref-66">66</xref>,<xref ref-type="bibr" rid="ref-48">48</xref>,<xref ref-type="bibr" rid="ref-89">89</xref>]</td>
<td>1990&#x2013;2016</td>
<td>Day</td>
<td>The dataset collects daily exchange rates from 1990 to 2016 for eight countries.</td>
</tr>
<tr>
<td>Gold prices</td>
<td>[<xref ref-type="bibr" rid="ref-94">94</xref>]</td>
<td>2014.1&#x2013;2018.4</td>
<td>Day</td>
<td>The dataset contains daily gold<break/>prices (U.S. dollars) from<break/>2014.1 to 2018.4.</td>
</tr>
<tr>
<td>Stock opening prices</td>
<td>[<xref ref-type="bibr" rid="ref-90">90</xref>]</td>
<td>2007&#x2013;2016</td>
<td>Day</td>
<td>The dataset collects daily opening prices for 50 stocks in 10 sectors in Financial Yahoo from 2007&#x2013;2016.</td>
</tr>
<tr>
<td>CRSP&#x2019;s stocks</td>
<td>[<xref ref-type="bibr" rid="ref-95">95</xref>]</td>
<td>&#x2013;</td>
<td>Day</td>
<td>The dataset is from CRSP and<break/>includes individual stock returns and prices, among other things.</td>
</tr>
<tr>
<td>Shanghai composite</td>
<td>[<xref ref-type="bibr" rid="ref-96">96</xref>]</td>
<td>2005.1&#x2013;2017.6</td>
<td>Day</td>
<td>This dataset records the daily SSE indices from 2005.01&#x2013;2017.06.</td>
</tr>
<tr>
<td>Finance Japan</td>
<td>[<xref ref-type="bibr" rid="ref-97">97</xref>]</td>
<td>2003.1&#x2013;2016.12</td>
<td>4 months</td>
<td>The dataset was collected by the Ministry of Finance of Japan and records general partnerships, limited partnerships, limited liability companies and joint stock companies from the first quarter of 2003 to the fourth quarter of 2016.</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-10">
<label>Table 10</label>
<caption>
<title>Meteorology dataset</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Ref.</th>
<th>Time range</th>
<th align="center">Time interval</th>
<th>Information</th>
</tr>
</thead>
<tbody>
<tr>
<td>Beijing<break/>PM2.5</td>
<td>[<xref ref-type="bibr" rid="ref-34">34</xref>],</td>
<td>2010.1.1&#x2013;2014.12.31</td>
<td>Hour</td>
<td>The dataset contains hourly PM2.5 data and associated meteorological data for Beijing, China.</td>
</tr>
<tr>
<td>WTH</td>
<td>[<xref ref-type="bibr" rid="ref-66">66</xref>]</td>
<td>2020</td>
<td>10 min</td>
<td>The dataset records weather conditions throughout 2020.</td>
</tr>
<tr>
<td>Hangzhou temperature</td>
<td>[<xref ref-type="bibr" rid="ref-96">96</xref>]</td>
<td>2011.1&#x2013;2017.1</td>
<td>Day</td>
<td>This dataset records the daily average temperature of Hangzhou from 2011.1 to 2017.1.</td>
</tr>
</tbody>
</table>
</table-wrap></p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Dividing the Dataset into Training, Validation, and Test Subsets</title>
<p>One of the most critical aspects of building forecasting models is the proper division of data into training, validation, and test sets. This division is essential for evaluating the model&#x2019;s ability to generalize to unseen data and ensuring that it does not overfit.
<list list-type="bullet">
<list-item>
<p><bold>Training Set:</bold> The training subset serves as the foundation for developing the forecasting model, enabling it to learn latent patterns and temporal dependencies within the data. For effective time series forecasting, the training period must be carefully selected to encompass a sufficiently extensive duration that captures all critical temporal characteristics.</p></list-item>
<list-item>
<p><bold>Validation Set:</bold> The validation set serves a critical role in model development by enabling hyperparameter optimization (e.g., network depth, learning rate schedules) and model selection. This intermediate dataset provides an unbiased performance assessment during iterative training, acting as an early stopping mechanism to mitigate overfitting while ensuring the model generalizes well to unseen data.</p></list-item>
<list-item>
<p><bold>Test Set:</bold> The test set serves as the gold standard for evaluating model performance, exclusively employed after completing all training and validation phases. This carefully withheld dataset simulates real-world deployment conditions by providing completely unseen data, enabling rigorous assessment of the model&#x2019;s generalization capacity.</p></list-item>
</list></p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Cross-Validation Techniques for Time Series</title>
<p>Traditional cross-validation techniques, such as k-fold cross-validation, are not always suitable for time series data due to the temporal dependencies. Instead, methods like TimeSeriesSplit or Walk-Forward Validation are preferred. These techniques involve using earlier data to predict later data while maintaining the temporal structure, ensuring that the validation process mimics the real-world forecasting scenario.
<list list-type="bullet">
<list-item>
<p><bold>TimeSeriesSplit:</bold> This method involves splitting the data into several folds, where each fold is used as a validation set in turn, while the training set grows progressively larger with each fold. This allows the model to be trained on more data while still being validated on unseen data.</p></list-item>
<list-item>
<p><bold>Walk-Forward Validation:</bold> In this method, the model is trained on the first portion of the time series, and then tested on the subsequent portion. The process is repeated by moving the training and testing windows forward in time.</p></list-item>
</list></p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Handling Data Imbalances and Outliers</title>
<p>In time series data, imbalances and outliers can often skew model performance. For instance, rare events (e.g., sudden stock market crashes or natural disasters) may disproportionately affect the dataset. Handling such events requires careful preprocessing, including the use of robust scaling, outlier detection, and data augmentation techniques. Moreover, these rare events should be treated with caution during validation to avoid misleading results.</p>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Data Augmentation for Time Series</title>
<p>To enhance model robustness and performance, especially when the dataset is small, data augmentation techniques can be employed. For time series forecasting, this might include methods like:
<list list-type="bullet">
<list-item>
<p><bold>Time Warping:</bold> Randomly stretching or compressing the time axis to generate new variations of the original time series.</p></list-item>
<list-item>
<p><bold>Jittering:</bold> Adding small random noise to the data to simulate different scenarios.</p></list-item>
<list-item>
<p><bold>Bootstrapping:</bold> Creating synthetic time series by resampling with replacement from the original data.</p></list-item>
</list></p>
<p>These techniques help to artificially enlarge the dataset, allowing the model to learn more diverse patterns and improve generalization.</p>
</sec>
<sec id="s4_6">
<label>4.6</label>
<title>Performance Evaluation Metrics for Time Series Forecasting (TSF)</title>
<p>In TSF, evaluation metrics serve as crucial tools. These metrics enable us to assess the forecasting capabilities of various models. By leveraging these metrics, we can objectively compare the performance of different models and identify the one that best addresses real-world problems. They provide a standardized approach to measuring a model&#x2019;s accuracy and precision, facilitating the selection of the most suitable model for predicting future time series data. The insights derived from these metrics allow us to choose the optimal model, thereby enhancing the effectiveness of practical applications. This section presents several widely used evaluation metrics for TSF tasks.
<list list-type="simple">
<list-item><label>1.</label><p><bold>Mean Square Error (MSE)</bold></p>
<p>MSE [<xref ref-type="bibr" rid="ref-98">98</xref>] is defined as the average of the squared differences between predicted and actual values. The formula is as follows:<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mi>M</mml:mi><mml:mi>S</mml:mi><mml:mi>E</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>m</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></disp-formula>where <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the true value, <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the predicted value, and <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:mi>m</mml:mi></mml:math></inline-formula> denotes the number of samples.</p></list-item>
<list-item><label>2.</label><p><bold>Root Mean Square Error (RMSE)</bold></p>
<p>RMSE [<xref ref-type="bibr" rid="ref-99">99</xref>] is the square root of the MSE. A smaller RMSE value reflects better predictive capabilities of the model [<xref ref-type="bibr" rid="ref-100">100</xref>]. The calculation formula for RMSE is:<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mi>R</mml:mi><mml:mi>M</mml:mi><mml:mi>S</mml:mi><mml:mi>E</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msqrt><mml:mfrac><mml:mn>1</mml:mn><mml:mi>m</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:msqrt></mml:math></disp-formula>where <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the true value, <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the predicted value, and <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mi>m</mml:mi></mml:math></inline-formula> denotes the number of samples.</p></list-item>
<list-item><label>3.</label><p><bold>Mean Absolute Error (MAE)</bold></p>
<p>MAE [<xref ref-type="bibr" rid="ref-101">101</xref>] is the average of the absolute differences between predicted and actual values. A smaller MAE value indicates enhanced predictive ability of the model [<xref ref-type="bibr" rid="ref-100">100</xref>]. The formula for MAE is presented below:<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:mi>M</mml:mi><mml:mi>A</mml:mi><mml:mi>E</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>m</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the true value, <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the predicted value, and <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mi>m</mml:mi></mml:math></inline-formula> denotes the number of samples.</p></list-item>
<list-item><label>4.</label><p><bold>Mean Absolute Percentage Error (MAPE)</bold></p>
<p>MAPE [<xref ref-type="bibr" rid="ref-102">102</xref>] is the average of the absolute percentage.<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:mi>M</mml:mi><mml:mi>A</mml:mi><mml:mi>P</mml:mi><mml:mi>E</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>m</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mo>|</mml:mo><mml:mfrac><mml:mrow><mml:mover><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:mo>|</mml:mo></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mn>100</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></disp-formula>where <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the true value, <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the predicted value, and <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mi>m</mml:mi></mml:math></inline-formula> denotes the number of samples.</p></list-item>
</list></p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion and Future Directions</title>
<p>This paper provides an in-depth exploration of the fundamental concepts and definitions related to time series forecasting. By categorizing TSF algorithms based on their underlying network structures, we have divided them into four main categories: CNN-based models, RNN-based models, GNN-based models, and Transformer-based models. We have elaborated on the concepts, principles, and applications of these models and reviewed the state-of-the-art approaches within each category. Additionally, we have presented commonly used datasets and performance evaluation metrics for TSF tasks.</p>
<p>Looking ahead, it is evident that most current TSF models are primarily designed for sequence data with equal time intervals and lack the capability to handle TSF tasks involving datasets with irregular time intervals. To address this limitation, interpolation, filtering, or other techniques could be integrated into the model architecture to effectively manage TSF problems with unequal time intervals. Furthermore, incorporating domain knowledge and considering external privacy factors into time series forecasting models can enhance the accuracy and interpretability of the predictions. These advancements hold significant potential for further improving the performance and applicability of TSF models in real-world scenarios.</p>
</sec>
</body>
<back>
<ack>
<p>The authors received no specific support for this study.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This research was funded by Natural Science Foundation of Heilongjiang Province, grant number LH2023F020.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualization, Gaoyong Lu and Wei Li; methodology, Yang Ou and Zhihong Wang; software, Zhihong Wang and Yingnan Qu; validation, Yingsheng Xia, Dibin Tang and Zhihong Wang; formal analysis, Igor Kotenko; investigation, Wei Li; resources, Gaoyong Lu; data curation, Yang Ou; writing&#x2014;original draft preparation, Zhihong Wang and Yingnan Qu; writing&#x2014;review and editing, Gaoyong Lu and Wei Li; visualization, Yingsheng Xia and Dibin Tang; supervision, Wei Li; project administration, Gaoyong Lu; funding acquisition, Wei Li. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>Data available on request from the authors.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kim</surname></string-name> <string-name> <given-names>KJ</given-names></string-name></person-group>. <article-title>Financial time series forecasting using support vector machines</article-title>. <source>Neurocomputing</source>. <year>2003</year>;<volume>55</volume>(<issue>1&#x2013;2</issue>):<fpage>307</fpage>&#x2013;<lpage>19</lpage>. doi:<pub-id pub-id-type="doi">10.1016/S0925-2312(03)00372-2</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kaastra</surname> <given-names>I</given-names></string-name>, <string-name><surname>Boyd</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Designing a neural network for forecasting financial and economic time series</article-title>. <source>Neurocomputing</source>. <year>1996</year>;<volume>10</volume>(<issue>3</issue>):<fpage>215</fpage>&#x2013;<lpage>36</lpage>. doi:<pub-id pub-id-type="doi">10.1016/0925-2312(95)00039-9</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Financial time series forecasting model based on CEEMDAN and LSTM</article-title>. <source>Phys A Stat Mech Appl</source>. <year>2019</year>;<volume>519</volume>:<fpage>127</fpage>&#x2013;<lpage>39</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.physa.2018.11.061</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>YF</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>B</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>L</given-names></string-name>, <string-name><surname>Thuraisingham</surname> <given-names>B</given-names></string-name>, <string-name><surname>Brandt</surname> <given-names>PT</given-names></string-name>, <string-name><surname>D&#x2019;Orazio</surname> <given-names>VJ</given-names></string-name></person-group>. <article-title>Data-driven time series forecasting for social studies using spatio-temporal graph neural networks</article-title>. In: <conf-name>Proceedings of the 2021 Conference on Information Technology for Social Good; 2021 Sep 9&#x2013;11; Rome, Italy</conf-name>. p. <fpage>61</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3462203.3475929</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Taylor</surname> <given-names>CJ</given-names></string-name>, <string-name><surname>Pedregal</surname> <given-names>DJ</given-names></string-name>, <string-name><surname>Young</surname> <given-names>PC</given-names></string-name>, <string-name><surname>Tych</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Environmental time series analysis and forecasting with the captain toolbox</article-title>. <source>Environ Modell Softw</source>. <year>2007</year>;<volume>22</volume>(<issue>6</issue>):<fpage>797</fpage>&#x2013;<lpage>814</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.envsoft.2006.03.002</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Murat</surname> <given-names>M</given-names></string-name>, <string-name><surname>Malinowska</surname> <given-names>I</given-names></string-name>, <string-name><surname>Gos</surname> <given-names>M</given-names></string-name>, <string-name><surname>Krzyszczak</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Forecasting daily meteorological time series using ARIMA and regression models</article-title>. <source>Int Agrophys</source>. <year>2018</year>;<volume>32</volume>(<issue>2</issue>):<fpage>253</fpage>&#x2013;<lpage>64</lpage>. doi:<pub-id pub-id-type="doi">10.1515/intag-2017-0007</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zheng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Traffic flow forecast through time series analysis based on deep learning</article-title>. <source>IEEE Access</source>. <year>2020</year>;<volume>8</volume>:<fpage>82562</fpage>&#x2013;<lpage>70</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2020.2990738</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lippi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Bertini</surname> <given-names>M</given-names></string-name>, <string-name><surname>Frasconi</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Short-term traffic flow forecasting: an experimental comparison of time-series analysis and supervised learning</article-title>. <source>IEEE Trans Intell Transp Syst</source>. <year>2013</year>;<volume>14</volume>(<issue>2</issue>):<fpage>871</fpage>&#x2013;<lpage>82</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TITS.2013.2247040</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>M</given-names></string-name>, <string-name><surname>Li</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Long sequence time-series forecasting with deep learning: a survey</article-title>. <source>Inf Fusion</source>. <year>2023</year>;<volume>97</volume>:<fpage>101819</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.inffus.2023.101819</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>N</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Time series multi-step forecasting based on memory network for the prognostics and health management in freight train braking system</article-title>. <source>IEEE Trans Intell Transp Syst</source>. <year>2023</year>;<volume>24</volume>(<issue>8</issue>):<fpage>8149</fpage>&#x2013;<lpage>62</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TITS.2023.3266227</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Noble</surname> <given-names>WS</given-names></string-name></person-group>. <article-title>What is a support vector machine? </article-title> <source>Nat Biotechnol</source>. <year>2006</year>;<volume>24</volume>(<issue>12</issue>):<fpage>1565</fpage>&#x2013;<lpage>7</lpage>.; <pub-id pub-id-type="pmid">17160063</pub-id></mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Eddy</surname> <given-names>SR</given-names></string-name></person-group>. <article-title>Hidden markov models</article-title>. <source>Curr Opin Struct Biol</source>. <year>1996</year>;<volume>6</volume>(<issue>3</issue>):<fpage>361</fpage>&#x2013;<lpage>5</lpage>. doi:<pub-id pub-id-type="doi">10.1016/S0959-440X(96)80056-X</pub-id>; <pub-id pub-id-type="pmid">8804822</pub-id></mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Devlin</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>MW</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>K</given-names></string-name>, <string-name><surname>Toutanova</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Bert: pre-training of deep bidirectional transformers for language understanding</article-title>. <comment>arXiv:1810.04805. 2018</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1810.04805</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hewamalage</surname> <given-names>H</given-names></string-name>, <string-name><surname>Bergmeir</surname> <given-names>C</given-names></string-name>, <string-name><surname>Bandara</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Recurrent neural networks for time series forecasting: current status and future directions</article-title>. <source>Int J Forecast</source>. <year>2021</year>;<volume>37</volume>(<issue>1</issue>):<fpage>388</fpage>&#x2013;<lpage>427</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ijforecast.2020.06.008</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Benidis</surname> <given-names>K</given-names></string-name>, <string-name><surname>Rangapuram</surname> <given-names>SS</given-names></string-name>, <string-name><surname>Flunkert</surname> <given-names>V</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Maddix</surname> <given-names>D</given-names></string-name>, <string-name><surname>Turkmen</surname> <given-names>C</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Deep learning for time series forecasting: tutorial and literature survey</article-title>. <source>ACM Comput Surv</source>. <year>2022</year>;<volume>55</volume>(<issue>6</issue>):<fpage>1</fpage>&#x2013;<lpage>36</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3533382</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Kim</surname> <given-names>J</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>H</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>H</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>D</given-names></string-name>, <string-name><surname>Yoon</surname> <given-names>S</given-names></string-name></person-group>. <article-title>A comprehensive survey of time series forecasting: architectural diversity and open challenges</article-title>. <comment>arXiv:2411.05793. 2024</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2411.05793</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Bai</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kolter</surname> <given-names>JZ</given-names></string-name>, <string-name><surname>Koltun</surname> <given-names>V</given-names></string-name></person-group>. <article-title>An empirical evaluation of generic convolutional and recurrent networks for sequence modeling</article-title>. <comment>arXiv:1803.01271. 2018</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1803.01271</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Ota</surname> <given-names>K</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Everything is image: CNN-based short-term electrical load forecasting for smart grid</article-title>. In: <conf-name>2017 14th International Symposium on Pervasive Systems, Algorithms and Networks &#x0026; 2017 11th International Conference on Frontier of Computer Science and Technology &#x0026; 2017 Third International Symposium of Creative Computing (ISPAN-FCST-ISCC); 2017 Jun 21&#x2013;23; Exeter, UK</conf-name>. p. <fpage>344</fpage>&#x2013;<lpage>51</lpage>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Borovykh</surname> <given-names>A</given-names></string-name>, <string-name><surname>Bohte</surname> <given-names>S</given-names></string-name>, <string-name><surname>Oosterlee</surname> <given-names>CW</given-names></string-name></person-group>. <article-title>Conditional time series forecasting with convolutional neural networks</article-title>. <comment>arXiv:1703.04691. 2017</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1703.04691</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Dong</surname> <given-names>X</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>L</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Short-term load forecasting in smart grid: a combined CNN and K-means clustering approach</article-title>. In: <conf-name>2017 IEEE International Conference on Big Data and Smart Computing (BigComp); 2017 Feb 13&#x2013;16; Jeju Island, Republic of Korea</conf-name>. p. <fpage>119</fpage>&#x2013;<lpage>25</lpage>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Cheng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Towards better forecasting by fusing near and distant future visions</article-title>. In: <conf-name>Proceedings of the 2020 AAAI Conference on Artificial Intelligence; 2020 Feb 7&#x2013;12; New York, NY, USA</conf-name>. p. <fpage>3593</fpage>&#x2013;<lpage>600</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v34i04.5766</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>M</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Lai</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>L</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Scinet: time series modeling and forecasting with sample convolution and interaction</article-title>. <source>Adv Neural Inf Process Syst</source>. <year>2022</year>;<volume>35</volume>:<fpage>5816</fpage>&#x2013;<lpage>28</lpage>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2106.09305</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Peng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>F</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xiao</surname></string-name> <string-name> <given-names>Y</given-names></string-name></person-group>. <article-title>Micn: multi-scale local and global context modeling for long-term series forecasting</article-title>. In: <conf-name>The Eleventh International Conference on Learning Representations; 2023 May 1&#x2013;5; Kigali, Rwanda</conf-name>. p. <fpage>1</fpage>&#x2013;<lpage>22</lpage>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Long</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Timesnet: temporal 2D-variation modeling for general time series analysis</article-title>. <comment>arXiv:2210.02186. 2022</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2210.02186</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lai</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Jensen</surname> <given-names>CS</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Lightcts: a lightweight framework for correlated time series forecasting</article-title>. <source>Proc ACM Manag Data</source>. <year>2023</year>;<volume>1</volume>(<issue>2</issue>):<fpage>125</fpage>. doi:<pub-id pub-id-type="doi">10.1145/3589270</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Luo</surname> <given-names>D</given-names></string-name>, <string-name><surname>Wang</surname></string-name> <string-name> <given-names>X</given-names></string-name></person-group>. <article-title>Cross-LKTCN: modern convolution utilizing cross-variable dependency for multivariate time series forecasting dependency for multivariate time series forecasting</article-title>. <comment>arXiv:2306.02326. 2023</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2306.02326</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Gong</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Patchmixer: a patch-mixing architecture for long-term time series forecasting</article-title>. <comment>arXiv:2310.00655. 2023</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2310.00655</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Luo</surname> <given-names>D</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Moderntcn: a modern pure convolution structure for general time series analysis</article-title>. In: <conf-name>The Twelfth International Conference on Learning Representations; 2024 May 7&#x2013;11; Vienna, Austria</conf-name>. p. <fpage>1</fpage>&#x2013;<lpage>43</lpage>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Elman</surname> <given-names>JL</given-names></string-name></person-group>. <article-title>Finding structure in time</article-title>. <source>Cogn Sci</source>. <year>1990</year>;<volume>14</volume>(<issue>2</issue>):<fpage>179</fpage>&#x2013;<lpage>211</lpage>. doi:<pub-id pub-id-type="doi">10.1016/0364-0213(90)90002-E</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Graves</surname> <given-names>A</given-names></string-name></person-group>. <source>Long short-term memory. Supervised sequence labelling with recurrent neural networks</source>. <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2012</year>. p. <fpage>37</fpage>&#x2013;<lpage>45</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-642-24797-2_4</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Cho</surname> <given-names>K</given-names></string-name>, <string-name><surname>Van Merri&#x00EB;nboer</surname> <given-names>B</given-names></string-name>, <string-name><surname>Gulcehre</surname> <given-names>C</given-names></string-name>, <string-name><surname>Bahdanau</surname> <given-names>D</given-names></string-name>, <string-name><surname>Bougares</surname> <given-names>F</given-names></string-name>, <string-name><surname>Schwenk</surname> <given-names>H</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Learning phrase representations using RNN encoder-decoder for statistical machine translation</article-title>. <comment>arXiv:1406.1078. 2014</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1406.1078</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lai</surname> <given-names>G</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>WC</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Modeling long-and short-term temporal patterns with deep neural networks</article-title>. In: <conf-name>The 41st International ACM SIGIR Conference on Research &#x0026; Development in Information Retrieval; 2018 Jul 8&#x2013;12; Ann Arbor, MI, USA</conf-name>. p. <fpage>95</fpage>&#x2013;<lpage>104</lpage>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1703.07015</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Qin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Song</surname> <given-names>D</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>W</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Cottrell</surname> <given-names>G</given-names></string-name></person-group>. <article-title>A dual-stage attention-based recurrent neural network for time series prediction</article-title>. <comment>arXiv:1704.02971. 2017</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1704.02971</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Chang</surname> <given-names>YY</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>FY</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>YH</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>SD</given-names></string-name></person-group>. <article-title>A memory-network based solution for multivariate time-series forecasting</article-title>. <comment>arXiv:1809.02105. 2018</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1809.021057</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Fan</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Yuan</surname> <given-names>R</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Multi-horizon time series forecasting with temporal attention learning</article-title>. In: <conf-name>Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery &#x0026; Data Mining; 2019 Aug 4&#x2013;8; Anchorage, AK, USA</conf-name>. p. <fpage>2527</fpage>&#x2013;<lpage>35</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3292500.3330662</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jung</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Jung</surname> <given-names>J</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>B</given-names></string-name>, <string-name><surname>Han</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Long short-term memory recurrent neural network for modeling temporal patterns in long-term power forecasting for solar PV facilities: case study of South Korea</article-title>. <source>J Clean Prod</source>. <year>2020</year>;<volume>250</volume>:<fpage>119476</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jclepro.2019.119476</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Du</surname> <given-names>S</given-names></string-name>, <string-name><surname>Li</surname> <given-names>T</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Horng</surname> <given-names>SJ</given-names></string-name></person-group>. <article-title>Multivariate time series forecasting via attention-based encoder-decoder framework</article-title>. <source>Neurocomputing</source>. <year>2020</year>;<volume>388</volume>:<fpage>269</fpage>&#x2013;<lpage>79</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neucom.2019.12.118</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Gangopadhyay</surname> <given-names>T</given-names></string-name>, <string-name><surname>Tan</surname> <given-names>SY</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Meng</surname> <given-names>R</given-names></string-name>, <string-name><surname>Sarkar</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Spatiotemporal attention for multivariate time series prediction and interpretation</article-title>. In: <conf-name>ICASSP 2021&#x2014;2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2021</year>. p. <fpage>3560</fpage>&#x2013;<lpage>4</lpage>. doi:<pub-id pub-id-type="doi">10.1109/icassp39728.2021.9413914</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Schirmer</surname> <given-names>M</given-names></string-name>, <string-name><surname>Eltayeb</surname> <given-names>M</given-names></string-name>, <string-name><surname>Lessmann</surname> <given-names>S</given-names></string-name>, <string-name><surname>Rudolph</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Modeling irregular time series with continuous recurrent units</article-title>. <comment>arXiv:2111.11344. 2022</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2111.11344</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Jia</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Hao</surname> <given-names>X</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Guo</surname></string-name> <string-name> <given-names>S</given-names></string-name>, <string-name><surname>Wan</surname></string-name> <string-name> <given-names>H</given-names></string-name></person-group>. <chapter-title>Witran: water-wave information transmission and recurrent acceleration network for long-range time series forecasting</chapter-title>. In: <source>Advances in neural information processing systems</source>. <publisher-loc>Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates, Inc.</publisher-name>; <year>2023</year>. p. <fpage>12389</fpage>&#x2013;<lpage>456</lpage>. doi:<pub-id pub-id-type="doi">10.5555/3666122.3666666</pub-id>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Veli&#x010D;kovi&#x0107;</surname> <given-names>P</given-names></string-name>, <string-name><surname>Cucurull</surname> <given-names>G</given-names></string-name>, <string-name><surname>Casanova</surname> <given-names>A</given-names></string-name>, <string-name><surname>Romero</surname> <given-names>A</given-names></string-name>, <string-name><surname>Li&#x00F2;</surname> <given-names>P</given-names></string-name>, <string-name><surname>Bengio</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Graph attention networks</article-title>. <comment>arXiv:1710.10903. 2018</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1710.10903</pub-id>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Kipf</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Semi-supervised classification with graph convolutional networks</article-title>. <comment>arXiv:1609.02907. 2016</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1609.02907</pub-id>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Yin</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Spatio-temporal graph convolutional networks: a deep learning framework for traffic forecasting</article-title>. In: <conf-name>Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence. International Joint Conferences on Artificial Intelligence Organization; 2018 Jul 13&#x2013;19; Stockholm, Sweden</conf-name>. p. <fpage>3634</fpage>&#x2013;<lpage>40</lpage>. doi:<pub-id pub-id-type="doi">10.24963/ijcai.2018/505</pub-id>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Guo</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>N</given-names></string-name>, <string-name><surname>Song</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wan</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Attention based spatial-temporal graph convolutional networks for traffic flow forecasting</article-title>. In: <conf-name>Proceedings of the 2019 AAAI Conference on Artificial Intelligence; 2019 Jan 27&#x2013;Feb 1; Honolulu, HI, USA</conf-name>. p. <fpage>922</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v33i01.3301922</pub-id>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Dynamic spatiotemporal interactive graph neural network for multivariate time series forecasting</article-title>. <source>Knowl Based Syst</source>. <year>2023</year>;<volume>280</volume>:<fpage>110995</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.knosys.2023.110995</pub-id>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Meng</surname> <given-names>G</given-names></string-name>, <string-name><surname>Xiang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Spatio-temporal graph structure learning for traffic forecasting</article-title>. In: <conf-name>Proceedings of the AAAI Conference on Artificial Intelligence; 2020 Feb 7&#x2013;12; New York, NY, USA</conf-name>. p. <fpage>1177</fpage>&#x2013;<lpage>85</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v34i01.5470</pub-id>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Yin</surname> <given-names>C</given-names></string-name>, <string-name><surname>Xiao</surname> <given-names>P</given-names></string-name>, <string-name><surname>Li</surname> <given-names>K</given-names></string-name>, <string-name><surname>Tan</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Attention-based spatial-temporal convolution gated recurrent unit for traffic flow forecasting</article-title>. <source>Entropy</source>. <year>2023</year>;<volume>25</volume>(<issue>6</issue>):<fpage>938</fpage>. doi:<pub-id pub-id-type="doi">10.3390/e25060938</pub-id>; <pub-id pub-id-type="pmid">37372282</pub-id></mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Long</surname> <given-names>G</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Connecting the dots: multivariate time series forecasting with graph neural networks</article-title>. In: <conf-name>Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery &#x0026; Data Mining; 2020 Jul 6&#x2013;10; Online</conf-name>. p. <fpage>753</fpage>&#x2013;<lpage>63</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3394486.3403118</pub-id>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ling</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>X</given-names></string-name></person-group>. <article-title>A multi-scale residual graph convolution network with hierarchical attention for predicting traffic flow in urban mobility</article-title>. <source>Complex Intell Syst</source>. <year>2024</year>;<volume>10</volume>(<issue>3</issue>):<fpage>3305</fpage>&#x2013;<lpage>17</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s40747-023-01324-9</pub-id>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Cao</surname> <given-names>D</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Duan</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>C</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Spectral temporal graph neural network for multivariate time-series forecasting</article-title>. <comment>arXiv:2103.07719. 2021</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2103.07719</pub-id>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Pan</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Ke</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>AutoSTG: neural architecture search for predictions of spatio-temporal graph</article-title>. In: <conf-name>WWW&#x2019;21: Proceedings of the Web Conference 2021; 2021 Apr 19&#x2013;23; Ljubljana, Slovenia</conf-name>. p. <fpage>1846</fpage>&#x2013;<lpage>55</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3442381.3449816</pub-id>.</mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Han</surname> <given-names>L</given-names></string-name>, <string-name><surname>Du</surname> <given-names>B</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>L</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lv</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Dynamic and multi-faceted spatio-temporal deep learning for traffic speed forecasting</article-title>. In: <conf-name>Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery &#x0026; Data Mining; 2021 Aug 14&#x2013;18; Singapore</conf-name>. p. <fpage>547</fpage>&#x2013;<lpage>55</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3447548.3467275</pub-id>.</mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>W</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>F</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Decoupled dynamic spatial-temporal graph neural network for traffic forecasting</article-title>. <source>Proc VLDB Endow</source>. <year>2022</year>;<volume>15</volume>(<issue>11</issue>):<fpage>2733</fpage>&#x2013;<lpage>46</lpage>. doi:<pub-id pub-id-type="doi">10.14778/3551793.3551827</pub-id>.</mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cheng</surname> <given-names>D</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>F</given-names></string-name>, <string-name><surname>Xiang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Financial time series forecasting with multi-modality graph neural network</article-title>. <source>Pattern Recognit</source>. <year>2022</year>;<volume>121</volume>:<fpage>108218</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.patcog.2021.108218</pub-id>.</mixed-citation></ref>
<ref id="ref-55"><label>[55]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Crisis event summary generative model based on hierarchical multimodal fusion</article-title>. <source>Pattern Recognit</source>. <year>2023</year>;<volume>144</volume>:<fpage>109890</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.patcog.2023.109890</pub-id>.</mixed-citation></ref>
<ref id="ref-56"><label>[56]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>JW</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Multivariate time-series forecasting with temporal polynomial graph neural networks</article-title>. <source>Adv Neural Inf Process Syst</source>. <year>2022</year>;<volume>35</volume>:<fpage>19414</fpage>&#x2013;<lpage>26</lpage>. doi:<pub-id pub-id-type="doi">10.5555/3600270.3601681</pub-id>.</mixed-citation></ref>
<ref id="ref-57"><label>[57]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Cai</surname> <given-names>W</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Msgnet: learning multi-scale inter-series correlations for multivariate time series forecasting</article-title>. In: <conf-name>Proceedings of the 2024 AAAI Conference on Artificial Intelligence; 2024 Feb 20&#x2013;27; Vancouver, BC, Canada</conf-name>. p. <fpage>11141</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v38i10.28991</pub-id>.</mixed-citation></ref>
<ref id="ref-58"><label>[58]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Vaswani</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shazeer</surname> <given-names>N</given-names></string-name>, <string-name><surname>Parmar</surname> <given-names>N</given-names></string-name>, <string-name><surname>Uszkoreit</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jones</surname> <given-names>L</given-names></string-name>, <string-name><surname>Gomez</surname> <given-names>AN</given-names></string-name>, <etal>et al</etal></person-group>. <chapter-title>Attention is all you need</chapter-title>. In: <source>Advances in neural information processing systems</source>. <publisher-loc>Cambridge, MA, USA</publisher-loc>: <publisher-name>MIT Press</publisher-name>; <year>2017</year>. <volume>Vol. 30</volume>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1706.03762</pub-id>.</mixed-citation></ref>
<ref id="ref-59"><label>[59]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>S</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>X</given-names></string-name>, <string-name><surname>Xuan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>X</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>W</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>YX</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting</article-title>. In: <conf-name>Proceedings of the 33rd International Conference on Neural Information Processing Systems; 2019 Dec 8&#x2013;14; Vancouver, BC, Canada</conf-name>. p. <fpage>5243</fpage>&#x2013;<lpage>53</lpage>. doi:<pub-id pub-id-type="doi">10.5555/3454287.3454758</pub-id>.</mixed-citation></ref>
<ref id="ref-60"><label>[60]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>R</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name></person-group>. <article-title>All you need is transformer: RTT prediction for TCP based on deep learning approach</article-title>. In: <conf-name> 2021 International Conference on Digital Society and Intelligent Systems (DSInS); 2021 Nov 19&#x2013;21; Chengdu, China</conf-name>. p. <fpage>348</fpage>&#x2013;<lpage>51</lpage>. doi:<pub-id pub-id-type="doi">10.1109/dsins54396.2021.9670591</pub-id>.</mixed-citation></ref>
<ref id="ref-61"><label>[61]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jin</surname> <given-names>K</given-names></string-name>, <string-name><surname>Wi</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>E</given-names></string-name>, <string-name><surname>Kang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>TrafficBERT: pre-trained model with large-scale data for long-range traffic flow forecasting</article-title>. <source>Expert Syst Appl</source>. <year>2021</year>;<volume>186</volume>:<fpage>115738</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2021.115738</pub-id>.</mixed-citation></ref>
<ref id="ref-62"><label>[62]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Peng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>H</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Informer: beyond efficient transformer for long sequence time-series forecasting</article-title>. In: <conf-name>Proceedings of the 35th AAAI Conference on Artificial Intelligence; 2021 Feb 2&#x2013;9; Online</conf-name>. p. <fpage>11106</fpage>&#x2013;<lpage>15</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v35i12.17325</pub-id>.</mixed-citation></ref>
<ref id="ref-63"><label>[63]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Banerjee</surname> <given-names>S</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>M</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Spatial-temporal synchronous graph transformer network (STSGT) for COVID-19 forecasting</article-title>. <source>Smart Health</source>. <year>2022</year>;<volume>26</volume>:<fpage>100348</fpage>. doi:<pub-id pub-id-type="doi">10.1109/TPAMI.2023.3293516</pub-id>.</mixed-citation></ref>
<ref id="ref-64"><label>[64]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lim</surname> <given-names>B</given-names></string-name>, <string-name><surname>Ar&#x0131;k</surname> <given-names>S&#x00D6;</given-names></string-name>, <string-name><surname>Loeff</surname> <given-names>N</given-names></string-name>, <string-name><surname>Pfister</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Temporal fusion transformers for interpretable multi-horizon time series forecasting</article-title>. <source>Int J Forecast</source>. <year>2021</year>;<volume>37</volume>(<issue>4</issue>):<fpage>1748</fpage>&#x2013;<lpage>64</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ijforecast.2021.03.012</pub-id>.</mixed-citation></ref>
<ref id="ref-65"><label>[65]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Koprinska</surname> <given-names>I</given-names></string-name>, <string-name><surname>Rana</surname> <given-names>M</given-names></string-name></person-group>. <article-title>SSDNet: state space decomposition neural network for time series forecasting</article-title>. In: <conf-name>2021 IEEE International Conference on Data Mining (ICDM); 2021 Dec 7&#x2013;10; Auckland, New Zealand</conf-name>. p. <fpage>370</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2112.10251</pub-id>.</mixed-citation></ref>
<ref id="ref-66"><label>[66]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Long</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Autoformer: decomposition transformers with auto-correlation for long-term series forecasting</article-title>. <source>Adv Neural Inf Process Syst</source>. <year>2021</year>;<volume>34</volume>:<fpage>22419</fpage>&#x2013;<lpage>30</lpage>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2106.13008</pub-id>.</mixed-citation></ref>
<ref id="ref-67"><label>[67]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Qi</surname> <given-names>X</given-names></string-name>, <string-name><surname>Hou</surname> <given-names>K</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ou</surname> <given-names>W</given-names></string-name></person-group>. <article-title>From known to unknown: knowledge-guided transformer for time-series sales forecasting in Alibaba</article-title>. <comment>arXiv:2109.08381. 2021</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2109.08381</pub-id>.</mixed-citation></ref>
<ref id="ref-68"><label>[68]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Liao</surname> <given-names>C</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>W</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>AX</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Pyraformer: low-complexity pyramidal attention for long-range time series modeling and forecasting</article-title>. In: <conf-name>International Conference on Learning Representations; 2022 Apr 25&#x2013;29; Online</conf-name>. p. <fpage>1</fpage>&#x2013;<lpage>20</lpage>.</mixed-citation></ref>
<ref id="ref-69"><label>[69]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ma</surname> <given-names>S</given-names></string-name>, <string-name><surname>Guan</surname> <given-names>S</given-names></string-name>, <string-name><surname>He</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Nie</surname> <given-names>J</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>M</given-names></string-name></person-group>. <article-title>TPAD: temporal-pattern-based neural network model for anomaly detection in multivariate time series</article-title>. <source>IEEE Sens J</source>. <year>2023</year>;<volume>23</volume>(<issue>24</issue>):<fpage>30668</fpage>&#x2013;<lpage>82</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JSEN.2023.3327138</pub-id>.</mixed-citation></ref>
<ref id="ref-70"><label>[70]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Bian</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wan</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>F</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>D</given-names></string-name>, <string-name><surname>Lyu</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Oil temperature prediction method based on deep learning and digital twins</article-title>. In: <conf-name>Asian Conference on Pattern Recognition</conf-name>. <publisher-loc>Cham, Swizterland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2023</year>. p. <fpage>174</fpage>&#x2013;<lpage>84</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-47665-5_15</pub-id>.</mixed-citation></ref>
<ref id="ref-71"><label>[71]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>T</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>L</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Fedformer: frequency enhanced decomposed transformer for long-term series forecasting</article-title>. In: <conf-name>International Conference on Machine Learning</conf-name>. <publisher-loc>Westminster, UK</publisher-loc>: <publisher-name>PMLR</publisher-name>; <year>2022</year>. p. <fpage>27268</fpage>&#x2013;<lpage>86</lpage>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2201.12740</pub-id>.</mixed-citation></ref>
<ref id="ref-72"><label>[72]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Cirstea</surname> <given-names>R</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>C</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Kieu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>X</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Triformer: triangular, variable-specific attentions for long sequence multivariate time series forecasting</article-title>. In: <conf-name>The Thirty-First International Joint Conference on Artificial Intelligence (IJCAI-22). International Joint Conference on Artificial Intelligence (IJCAI); 2022 Jul 23&#x2013;29; Vienna, Austria</conf-name>. p. <fpage>1994</fpage>&#x2013;<lpage>2001</lpage>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2204.13767</pub-id>.</mixed-citation></ref>
<ref id="ref-73"><label>[73]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>W</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Peng</surname> <given-names>B</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>T</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Learning to rotate: quaternion transformer for complicated periodical time series forecasting</article-title>. In: <conf-name>Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining; 2022 Aug 14&#x2013;18; Washington, DC, USA</conf-name>. p. <fpage>146</fpage>&#x2013;<lpage>56</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3534678.3539234</pub-id>.</mixed-citation></ref>
<ref id="ref-74"><label>[74]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Liang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ke</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Airformer: predicting nationwide air quality in china with transformers</article-title>. In: <conf-name>Proceedings of the 37th AAAI Conference on Artificial Intelligence; 2023 Feb 7&#x2013;14; Washington, DC, USA</conf-name>. p. <fpage>14329</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v37i12.26676</pub-id>.</mixed-citation></ref>
<ref id="ref-75"><label>[75]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Nie</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Nguyen</surname> <given-names>NH</given-names></string-name>, <string-name><surname>Sinthong</surname> <given-names>P</given-names></string-name>, <string-name><surname>Kalagnanam</surname> <given-names>J</given-names></string-name></person-group>. <article-title>A time series is worth 64 words: long-term forecasting with transformers</article-title>. <comment>arXiv:2211.14730. 2022</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2211.14730</pub-id>.</mixed-citation></ref>
<ref id="ref-76"><label>[76]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Meng</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>G</given-names></string-name>, <string-name><surname>Wang</surname></string-name> <string-name> <given-names>H</given-names></string-name></person-group>. <article-title>Detformer: detect the reliable attention index for ultra-long time series forecasting</article-title>. In: <conf-name>International Conference on Intelligent Computing</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2023</year>. p. <fpage>457</fpage>&#x2013;<lpage>68</lpage>. doi:<pub-id pub-id-type="doi">10.1515/intag-2017-0007</pub-id>.</mixed-citation></ref>
<ref id="ref-77"><label>[77]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Shabani</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Abdi</surname> <given-names>AH</given-names></string-name>, <string-name><surname>Meng</surname> <given-names>L</given-names></string-name>, <string-name><surname>Sylvain</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Scaleformer: iterative multi-scale refining transformers for time series forecasting</article-title>. <comment>arXiv:2206.04038. 2023</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2206.04038</pub-id>.</mixed-citation></ref>
<ref id="ref-78"><label>[78]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>H</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Su</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>B</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Towards long-term time-series forecasting: feature, pattern, and distribution</article-title>. In: <conf-name>2023 IEEE 39th International Conference on Data Engineering (ICDE); 2023 Apr 3&#x2013;7; Anaheim, CA, USA</conf-name>. p. <fpage>1611</fpage>&#x2013;<lpage>24</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ICDE55515.2023.00127</pub-id>.</mixed-citation></ref>
<ref id="ref-79"><label>[79]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Jiang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Han</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>WX</given-names></string-name>, <string-name><surname>Wang</surname></string-name> <string-name> <given-names>J</given-names></string-name></person-group>. <article-title>Pdformer: propagation delay-aware dynamic long-range transformer for traffic flow prediction</article-title>. In: <conf-name>Proceedings of the AAAI Conference on Artificial Intelligence; 2023 Feb 7&#x2013;14; Washington, DC, USA</conf-name>. p. <fpage>4365</fpage>&#x2013;<lpage>73</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v37i4.25556</pub-id>.</mixed-citation></ref>
<ref id="ref-80"><label>[80]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>T</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>M</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>L</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>GCformer: an efficient solution for accurate and scalable long-term multivariate time series forecasting</article-title>. In: <conf-name>Proceedings of the 32nd ACM International Conference on Information and Knowledge Management; 2023 Oct 21&#x2013;25; Birmingham, UK</conf-name>. p. <fpage>3464</fpage>&#x2013;<lpage>73</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3583780.361513</pub-id>.</mixed-citation></ref>
<ref id="ref-81"><label>[81]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Meng</surname> <given-names>L</given-names></string-name>, <string-name><surname>Gu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>SageFormer: series-aware framework for long-term multivariate time-series forecasting</article-title>. <source>IEEE Internet Things J</source>. <year>2024</year>;<volume>11</volume>(<issue>10</issue>):<fpage>18435</fpage>&#x2013;<lpage>48</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JIOT.2024.3363451</pub-id>.</mixed-citation></ref>
<ref id="ref-82"><label>[82]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>B</given-names></string-name>, <string-name><surname>Cui</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Tsang</surname> <given-names>IW</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>DifFormer: multi-resolutional differencing transformer with dynamic ranging for time series analysis</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2023</year>;<volume>45</volume>(<issue>11</issue>):<fpage>13586</fpage>&#x2013;<lpage>98</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TPAMI.2023.3293516</pub-id>; <pub-id pub-id-type="pmid">37428671</pub-id></mixed-citation></ref>
<ref id="ref-83"><label>[83]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>F</given-names></string-name>, <string-name><surname>Shao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Dsformer: a double sampling transformer for multivariate time series long-term prediction</article-title>. In: <conf-name>Proceedings of the 32nd ACM International Conference on Information and Knowledge Management; 2023 Oct 21&#x2013;25; Birmingham, UK</conf-name>. p. <fpage>3062</fpage>&#x2013;<lpage>72</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3583780.361485</pub-id>.</mixed-citation></ref>
<ref id="ref-84"><label>[84]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>L</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>iTransformer: inverted transformers are effective for time series forecasting</article-title>. <comment>arXiv:2310.06625. 2023</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2310.06625</pub-id>.</mixed-citation></ref>
<ref id="ref-85"><label>[85]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>T</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Chen</surname></string-name><string-name> <given-names>H</given-names></string-name>, <string-name><surname>Huang</surname></string-name> <string-name> <given-names>S</given-names></string-name></person-group>. <article-title>Master: market-guided stock transformer for stock price forecasting</article-title>. In: <conf-name>Proceedings of the AAAI Conference on Artificial Intelligence; 2024 Feb 20&#x2013;27; Vancouver, BC, Canada</conf-name>. p. <fpage>162</fpage>&#x2013;<lpage>70</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v38i1.27767</pub-id>.</mixed-citation></ref>
<ref id="ref-86"><label>[86]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>B</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>R</given-names></string-name></person-group>. <article-title>CARD: channel aligned robust blend transformer for time series forecasting</article-title>. <comment>arXiv:2305.12095. 2024</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2305.12095</pub-id>.</mixed-citation></ref>
<ref id="ref-87"><label>[87]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>K</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Fang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>W</given-names></string-name>, <string-name><surname>Li</surname> <given-names>D</given-names></string-name></person-group>. <article-title>ContiFormer: continuous-time transformer for irregular time series modeling</article-title>. In: <conf-name>Proceedings of the 37th International Conference on Neural Information Processing Systems; 2023 Dec 10&#x2013;16; New Orleans, LA, USA</conf-name>. p. <fpage>47143</fpage>&#x2013;<lpage>75</lpage>. doi:<pub-id pub-id-type="doi">10.5555/3666122.3668164</pub-id>.</mixed-citation></ref>
<ref id="ref-88"><label>[88]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ni</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>W</given-names></string-name></person-group>. <article-title>BasisFormer: attention-based time series forecasting with learnable and interpretable basis</article-title>. In: <conf-name>Proceedings of the 37th International Conference on Neural Information Processing Systems; 2023 Dec 10&#x2013;16; New Orleans, LA, USA</conf-name>. p. <fpage>71222</fpage>&#x2013;<lpage>41</lpage>. doi:<pub-id pub-id-type="doi">10.5555/3666122.3669240</pub-id>.</mixed-citation></ref>
<ref id="ref-89"><label>[89]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yoo</surname> <given-names>J</given-names></string-name>, <string-name><surname>Kang</surname> <given-names>U</given-names></string-name></person-group>. <article-title>Attention-based autoregression for accurate and efficient multivariate time series forecasting</article-title>. In: <conf-name>Proceedings of the 2021 SIAM International Conference on Data Mining (SDM); 2021 Apr 29&#x2013;May 1; Alexandria, VA, USA</conf-name>. p. <fpage>531</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1137/1.9781611976700.60</pub-id>.</mixed-citation></ref>
<ref id="ref-90"><label>[90]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yao</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Forecasting wavelet transformed time series with attentive neural networks</article-title>. In: <conf-name>2018 IEEE International Conference on Data Mining (ICDM); 2018 Nov 17&#x2013;20; Singapore</conf-name>. p. <fpage>1452</fpage>&#x2013;<lpage>7</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ICDM.2018.002017</pub-id>.</mixed-citation></ref>
<ref id="ref-91"><label>[91]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Xiao</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>P</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Adversarial sparse transformer for time series forecasting</article-title>. <source>Adv Neural Inf Process Syst</source>. <year>2020</year>;<volume>33</volume>:<fpage>17105</fpage>&#x2013;<lpage>15</lpage>. doi:<pub-id pub-id-type="doi">10.5555/3495724.3497159</pub-id>.</mixed-citation></ref>
<ref id="ref-92"><label>[92]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Koprinska</surname> <given-names>I</given-names></string-name>, <string-name><surname>Rana</surname> <given-names>M</given-names></string-name></person-group>. <article-title>SpringNet: transformer and spring DTW for time series forecasting</article-title>. In: <conf-name>International Conference on Neural Information Processing</conf-name>. <publisher-loc>Cham, Swizterland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2020</year>. p. <fpage>616</fpage>&#x2013;<lpage>28</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-030-63836-8_51</pub-id>.</mixed-citation></ref>
<ref id="ref-93"><label>[93]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>W</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Mei</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Adaptive temporal-frequency network for time-series forecasting</article-title>. <source>IEEE Trans Knowl Data Eng</source>. <year>2020</year>;<volume>34</volume>(<issue>4</issue>):<fpage>1576</fpage>&#x2013;<lpage>87</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TKDE.2020.3003420</pub-id>.</mixed-citation></ref>
<ref id="ref-94"><label>[94]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Livieris</surname> <given-names>IE</given-names></string-name>, <string-name><surname>Pintelas</surname> <given-names>E</given-names></string-name>, <string-name><surname>Pintelas</surname> <given-names>P</given-names></string-name></person-group>. <article-title>A CNN-LSTM model for gold price time-series forecasting</article-title>. <source>Neural Comput Appl</source>. <year>2020</year>;<volume>32</volume>:<fpage>17351</fpage>&#x2013;<lpage>60</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00521-020-04867-x</pub-id>.</mixed-citation></ref>
<ref id="ref-95"><label>[95]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hou</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>An enriched time-series forecasting framework for long-short portfolio strategy</article-title>. <source>IEEE Access</source>. <year>2020</year>;<volume>8</volume>:<fpage>31992</fpage>&#x2013;<lpage>2002</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2020.2973037</pub-id>.</mixed-citation></ref>
<ref id="ref-96"><label>[96]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xiao</surname> <given-names>G</given-names></string-name></person-group>. <article-title>A novel time series forecasting model with deep learning</article-title>. <source>Neurocomputing</source>. <year>2020</year>;<volume>396</volume>:<fpage>302</fpage>&#x2013;<lpage>13</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neucom.2018.12.084</pub-id>.</mixed-citation></ref>
<ref id="ref-97"><label>[97]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yoshimi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Eguchi</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Forecasting corporate financial time series using multi-phase attention recurrent neural networks</article-title>. In: <conf-name>EDBT/ICDT Workshops</conf-name> <comment>2020 [Internet]. [cited 2025 Aug 6]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://star.informatik.rwth-aachen.de/Publications/CEUR-WS/Vol-2578/DARLIAP12.pdf">http://star.informatik.rwth-aachen.de/Publications/CEUR-WS/Vol-2578/DARLIAP12.pdf</ext-link>.</mixed-citation></ref>
<ref id="ref-98"><label>[98]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Bovik</surname> <given-names>AC</given-names></string-name></person-group>. <article-title>Mean squared error: love it or leave it? A new look at signal fidelity measures</article-title>. <source>IEEE Signal Process Mag</source>. <year>2009</year>;<volume>26</volume>(<issue>1</issue>):<fpage>98</fpage>&#x2013;<lpage>117</lpage>. doi:<pub-id pub-id-type="doi">10.1515/intag-2017-0007</pub-id>.</mixed-citation></ref>
<ref id="ref-99"><label>[99]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chai</surname> <given-names>T</given-names></string-name>, <string-name><surname>Draxler</surname> <given-names>RR</given-names></string-name></person-group>. <article-title>Root mean square error (RMSE) or mean absolute error (MAE)?&#x2014;Arguments against avoiding RMSE in the literature</article-title>. <source>Geosci Model Dev</source>. <year>2014</year>;<volume>7</volume>(<issue>3</issue>):<fpage>1247</fpage>&#x2013;<lpage>50</lpage>. doi:<pub-id pub-id-type="doi">10.5194/gmd-7-1247-2014</pub-id>.</mixed-citation></ref>
<ref id="ref-100"><label>[100]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>L</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Stock price prediction with attentive temporal convolution-based generative adversarial network</article-title>. <source>Array</source>. <year>2025</year>;<volume>25</volume>:<fpage>100374</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.array.2025.100374</pub-id>.</mixed-citation></ref>
<ref id="ref-101"><label>[101]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Willmott</surname> <given-names>CJ</given-names></string-name>, <string-name><surname>Matsuura</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Advantages of the mean absolute error (MAE) over the root mean square error (RMSE) in assessing average model performance</article-title>. <source>Clim Res</source>. <year>2005</year>;<volume>30</volume>(<issue>1</issue>):<fpage>79</fpage>&#x2013;<lpage>82</lpage>. doi:<pub-id pub-id-type="doi">10.3354/cr030079</pub-id>.</mixed-citation></ref>
<ref id="ref-102"><label>[102]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>De Myttenaere</surname> <given-names>A</given-names></string-name>, <string-name><surname>Golden</surname> <given-names>B</given-names></string-name>, <string-name><surname>Le Grand</surname> <given-names>B</given-names></string-name>, <string-name><surname>Rossi</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Mean absolute percentage error for regression models</article-title>. <source>Neurocomputing</source>. <year>2016</year>;<volume>192</volume>:<fpage>38</fpage>&#x2013;<lpage>48</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neucom.2015.12.114</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>