A Survey of Automated Time Series Forecasting: From Deep Learning to Foundation Models with Financial Applications

Main Article Content

Zhuoran Yang

Keywords

time series forecasting, machine learning, deep learning, transformer, foundation models, intelligent agents

Abstract

Time-series forecasting is a fundamental research topic in artificial intelligence and machine learning, with broad applications in finance, healthcare, energy management, transportation, and industrial systems. As the scale and complexity of temporal data continue to increase, forecasting methodologies have evolved significantly from traditional statistical models to modern intelligent forecasting systems. Understanding this technological evolution is essential for both researchers and practitioners seekin g to develop more accurate, scalable, and automated forecasting solutions. This survey reviews the development of automated time-series forecasting methods from a historical and technological perspective. The evolution of forecasting paradigms is examined across multiple stages, including statistical learning methods, machine learning approaches, deep learning architectures, Transformer-based forecasting models, AutoML frameworks, foundation models, and large-language-model-based intelligent agents. Particular attention is given to how these paradigms differ in terms of automation capability, long-term forecasting performance, interpretability, and application adaptability. The survey highlights the key motivations driving major technological transitions, including limitations in statistical assumptions, challenges in feature engineering, difficulties in long-sequence modeling, and the growing demand for generalization and autonomous decision-making capabilities. In addition, representative forecasting applications in the financial domain are reviewed to illustrate the practical impact of modern forecasting technologies. Finally, current challenges and future research directions are discussed, including multimodal forecasting, foundation model development, intelligent agen t systems, explainable forecasting, and computational efficiency. The findings suggest that time-series forecasting is gradually evolving from task-specific prediction models toward intelligent forecasting ecosystems capable of learning, reasoning, and autonomous decision-making. 

Abstract 11 | PDF Downloads 6

References

  • [1] R. J. Hyndman and G. Athanasopoulos, Forecasting: Principles and Practice, 3rd ed. Melbourne, Australia: OTexts, 2021. Available: https://otexts.com/fpp3/
  • [2] J. D. Hamilton, Time Series Analysis. Princeton, NJ, USA: Princeton University Press, 1994.
  • [3] O. B. Sezer, M. U. Gudelek, and A. M. Ozbayoglu, “Financial time series forecasting with deep learning: A systematic literature review: 2005–2019,” Applied Soft Computing, vol. 90, Art. no. 106181, 2020, doi: 10.1016/j.asoc.2020.106181.
  • [4] S. Makridakis, E. Spiliotis, and V. Assimakopoulos, “Statistical and machine learning forecasting methods: Concerns and ways forward,” PLOS ONE, vol. 13, no. 3, Art. no. e0194889, 2018, doi: 10.1371/journal.pone.0194889.
  • [5] B. Lim and S. Zohren, “Time-series forecasting with deep learning: A survey,” Philosophical Transactions of the Royal Society A, vol. 379, no. 2194, Art. no. 20200209, 2021, doi: 10.1098/rsta.2020.0209.
  • [6] G. E. P. Box and G. M. Jenkins, Time Series Analysis: Forecasting and Control, 2nd ed. San Francisco, CA, USA: Holden-Day, 1976.
  • [7] R. F. Engle, “Autoregressive conditional heteroscedasticity with estimates of the variance of United Kingdom inflation,” Econometrica, vol. 50, no. 4, pp. 987–1007, 1982, doi: 10.2307/1912773.
  • [8] T. Bollerslev, “Generalized autoregressive conditional heteroskedasticity,” Journal of Econometrics, vol. 31, no. 3, pp. 307–327, 1986, doi: 10.1016/0304-4076(86)90063-1.
  • [9] C. A. Sims, “Macroeconomics and reality,” Econometrica, vol. 48, no. 1, pp. 1–48, 1980, doi: 10.2307/1912017.
  • [10] C. Cortes and V. Vapnik, “Support-vector networks,” Machine Learning, vol. 20, no. 3, pp. 273–297, 1995, doi: 10.1007/BF00994018.
  • [11] L. Breiman, “Random forests,” Machine Learning, vol. 45, no. 1, pp. 5–32, 2001, doi: 10.1023/A:1010933404324.
  • [12] T. Chen and C. Guestrin, “XGBoost: A scalable tree boosting system,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 2016, pp. 785–794, doi: 10.1145/2939672.2939785.
  • [13] I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning. Cambridge, MA, USA: MIT Press, 2016.
  • [14] S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation, vol. 9, no. 8, pp. 1735–1780, 1997, doi: 10.1162/neco.1997.9.8.1735.
  • [15] K. Cho, B. van Merrië nboer, Ç. Gü lç ehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using RNN encoder–decoder for statistical machine translation,” in Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, Doha, Qatar, 2014, pp. 1724–1734.
  • [16] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems, vol. 30, 2017.
  • [17] Q. Wen, T. Zhou, C. Zhang, W. Chen, Z. Ma, J. Yan, and L. Sun, “Transformers in time series: A survey,” in Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, 2023, pp. 6778–6786, doi: 10.24963/ijcai.2023/759.
  • [18] H. Zhou, S. Zhang, J. Peng, S. Zhang, J. Li, H. Xiong, and W. Zhang, “Informer: Beyond efficient transformer for long sequence time-series forecasting,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, no. 12, 2021, pp. 11106–11115, doi: 10.1609/aaai.v35i12.17325.
  • [19] H. Wu, J. Xu, J. Wang, and M. Long, “Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting,” in Advances in Neural Information Processing Systems, vol. 34, 2021.
  • [20] R. Bommasani et al., “On the opportunities and risks of foundation models,” arXiv preprint arXiv:2108.07258, 2021.
  • [21] S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, “ReAct: Synergizing reasoning and acting in language models,” in Proceedings of the 11th International Conference on Learning Representations, 2023.
  • [22] T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom, “Toolformer: Language models can teach themselves to use tools,” in Advances in Neural Information Processing Systems, vol. 36, 2023.
  • [23] L. Wang et al., “A survey on large language model based autonomous agents,” Frontiers of Computer Science, vol. 18, Art. no. 186345, 2024, doi: 10.1007/s11704-024-40231-1.
  • [24] S. Makridakis, E. Spiliotis, and V. Assimakopoulos, “The M4 Competition: Results, findings, conclusion and way forward,” International Journal of Forecasting, vol. 34, no. 4, pp. 802–808, 2018, doi: 10.1016/j.ijforecast.2018.06.001.
  • [25] Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature, vol. 521, no. 7553, pp. 436–444, 2015, doi: 10.1038/nature14539.
  • [26] T. Zhou, Z. Ma, Q. Wen, X. Wang, L. Sun, and R. Jin, “FEDformer: Frequency enhanced decomposed transformer for long-term series forecasting,” in Proceedings of the 39th International Conference on Machine Learning, PMLR, vol. 162, 2022, pp. 27268–27286.
  • [27] Y. Nie, N. H. Nguyen, P. Sinthong, and J. Kalagnanam, “A time series is worth 64 words: Long-term forecasting with transformers,” in Proceedings of the International Conference on Learning Representations, 2023.
  • [28] A. F. Ansari et al., “Chronos: Learning the language of time series,” arXiv preprint arXiv:2403.07815, 2024.
  • [29] A. Garza, C. Challu, and M. Mergenthaler-Canseco, “TimeGPT-1,” arXiv preprint arXiv:2310.03589, 2023.
  • [30] A. Das, W. Kong, R. Sen, and Y. Zhou, “A decoder-only foundation model for time-series forecasting,” in Proceedings of the 41st International Conference on Machine Learning, PMLR, vol. 235, 2024, pp. 10148–10167.
  • [31] K. Rasul et al., “Lag-Llama: Towards foundation models for probabilistic time series forecasting,” arXiv preprint arXiv:2310.08278, 2024.
  • [32] F. Hutter, L. Kotthoff, and J. Vanschoren, Eds., Automated Machine Learning: Methods, Systems, Challenges. Cham, Switzerland: Springer, 2019, doi: 10.1007/978-3-030-05318-5.
  • [33] J. Bergstra and Y. Bengio, “Random search for hyper-parameter optimization,” Journal of Machine Learning Research, vol. 13, pp. 281–305, 2012.
  • [34] S. Y. Shah et al., “AutoAI-TS: AutoAI for time series forecasting,” arXiv preprint arXiv:2102.12347, 2021.
  • [35] T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, “Large language models are zero-shot reasoners,” in Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 22199–22213.
  • [36] X. Jin, Y. Park, D. C. Maddix, H. Wang, and Y. Wang, “Domain adaptation for time series forecasting via attention sharing,” in Proceedings of the 39th International Conference on Machine Learning, PMLR, vol. 162, 2022, pp. 10280–10297.
  • [37] Z. C. Lipton, “The mythos of model interpretability,” Queue, vol. 16, no. 3, pp. 31–57, 2018, doi: 10.1145/3236386.3241340. Also available as arXiv:1606.03490.
  • [38] Y. Gal and Z. Ghahramani, “Dropout as a Bayesian approximation: Representing model uncertainty in deep learning,” in Proceedings of the 33rd International Conference on Machine Learning, PMLR, vol. 48, 2016, pp. 1050–1059.
  • [39] T. Baltrušaitis, C. Ahuja, and L. -P. Morency, “Multimodal machine learning: A survey and taxonomy,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 41, no. 2, pp. 423–443, 2019, doi: 10.1109/TPAMI.2018.2798607.