Cross-Sectional Return Prediction in China’s A-Share Market Based on Lasso and XGBoost: A Comparison with the Fama–MacBeth Multi-Factor Model

Main Article Content

Xichun Zhang

Keywords

cross-sectional return prediction, Fama–MacBeth, Lasso, XGBoost, SHAP, out-of-sample OOS- R²

Abstract

In response to the problem of the widespread use of mixed prediction performance metrics in machine learning stock selection literature, the lack of empirical tests on the applicability boundaries of linear and tree-based models under low-dimensional factor settings, and the disconnection of most high-dimensional factor empirical conclusions from the actual scenarios of small and medium-sized quantitative institutions, this paper selects the monthly panel data of A-shares from 2013 to 2023 to avoid the leakage of look-ahead information, and builds a unified rolling-window forecasting framework of FM, Lasso-FM, and XGBoost. The model hyperparameters are selected based on the existing literature to ensure a fair comparison among the three models. It relies on OOS-R², IC, Rank-IC and long-short portfolio Sharpe ratio to carry out multi-dimensional evaluation, and carries out analysis in combination with the heterogeneity of market value grouping and SHAP decomposition. In addition, multiple sets of robustness checks are conducted. The empirical evidence shows that under the sample setting of the five basic factors in this article, the mean out-of-sample OOS- R² values of all three models are negative, the absolute return prediction performance is weak, and there is no significant statistical superiority or inferiority between the models; However, there is a deviation in the evaluation index, and XGBoost performs better in cross-sectional stock ranking and long–short portfolio performance, which is rooted in the fact that OOS-R² focuses on absolute error, and IC only focuses on the relative return strength of stocks; Based on the observation of this sample, it is difficult for XGBoost to release nonlinear modeling capability, and greater model flexibility has the possibility of inducing out-of-sample overfitting. The forecast stability of large-cap stocks is significantly higher than that of small-cap stocks, and the market value and ROA are important prediction features of the model. This paper explains the intrinsic mechanism of multi-indicator deviation, defines the applicable boundaries of factor dimensionality of different models, and provides a reference for target lightweight modeling for small and medium-sized institutions.

Abstract 0 | PDF Downloads 0