A Comparative Study of Traditional Statistical Models and Machine Learning Algorithms in Credit Risk Assessment

Main Article Content

Hanrun Jin

Keywords

credit risk assessment, traditional statistical models, machine learning, logistic regression, tree-based models, model performance evaluation

Abstract

Credit risk assessment underpins lending decisions, pricing strategies, portfolio management, and regulatory capital allocation within modern financial systems. Logistic regression has historically served as the dominant modeling framework in credit scoring due to its probabilistic coherence and interpretability. In recent years, advances in machine learning—particularly tree-based ensemble methods such as Random Forest and Gradient Boosting—have demonstrated strong predictive performance and often outperform traditional approaches in discrimination metrics such as the area under the ROC curve (AUC). However, the adoption of machine learning in credit risk modeling remains debated due to concerns regarding probability calibration, temporal robustness, interpretability, and regulatory governance. This paper provides a comprehensive comparison of traditional statistical models and tree-based machine learning approaches in credit risk assessment. Rather than focusing exclusively on discriminatory performance, the analysis adopts a multidimensional evaluation framework incorporating calibration quality and temporal stability. Drawing on foundational theory and recent empirical evidence, the paper argues that model adequacy in credit risk is inherently context dependent. A three-pillar framework—discrimination, calibration, and temporal robustness—is proposed to guide academic research and practical model deployment. The findings suggest that superior ranking performance does not necessarily imply superior decision quality and that effective credit risk modeling requires balancing predictive flexibility with probabilistic reliability and governance stability.

Abstract 20 | PDF Downloads 11

References

  • [1]Bolton, C. (2009). Logistic regression and its application in credit scoring. University of London.
  • [2]Brown, K., & Moles, P. (2008). Credit risk management. Edinburgh Business School, Heriot-Watt University.
  • [3]Xu, T. (2024). Comparative analysis of machine learning algorithms for consumer credit risk assessment. International Conference on Computational Intelligence and Applications (ICCIA).
  • [4]Vakrani, D. S., Padhye, P. S., Gupta, S. K., Roy, J. K., Nerlekar, V. S., & Parashar, N. (2026). Evaluating AI-driven credit scoring models versus traditional statistical techniques. Discover Artificial Intelligence, 6, 72. https://doi.org/10.1007/s44163-025-00772-1
  • [5]Oliveira, N. A. de, & Basso, L. F. C. (2025). Explaining corporate ratings transitions and defaults through machine learning. Algorithms, 18(10), 608. https://doi.org/10.3390/a18100608
  • [6]World Bank & International Committee on Credit Reporting. (2019). Credit scoring approaches guidelines. World Bank.
  • [7]Machado, M. R., Chen, D. T., & Osterrieder, J. R. (2025). An analytical approach to credit risk assessment using machine learning models. Decision Analytics Journal, 16, 100605. https://doi.org/10.1016/j.dajour.2025.100605