A Review of Distributed Computing Technologies for Financial Data Preprocessing
Main Article Content
Keywords
data preprocessing, financial data, distributed computing, Apache Flink
Abstract
Financial data preprocessing is a key link in financial analysis and modeling. With the exponential growth of data scale, the single-machine architecture is facing severe bottlenecks. Distributed computing provides a feasible path to break through performa nce limitations through multi-node collaborative processing. This article systematically sorts out the research status and technical progress of distributed computing in the field of financial data preprocessing. First of all, analyze the common characteri stics and pre-processing task genealogy of financial data, and combine Tang Yao’s stock linkage effect research to show the typical process of single-machine preprocessing pipelines and its scale bottlenecks; Then review the mainstream frameworks such as Hadoop, Spark and Flink from the perspective of technological evolution, and summ arize the three core empowerment mechanisms of data parallelism, computing parallelism and stream processing; On this basis, classify and summarize the application research of distributed computing in data cleaning, feature engineering, real-time processing and other links, and take Ma Chiyu’s financial news sentiment analysis based on SparkR as a case to verify the performance advantages of distributed preprocessing in actual tasks; Finally, we will comment on the limitations of existing research and look forward to the future direction of ad aptive preprocessing, explainable attribution, privacy protection calculation, etc. This article aims to provide a systematic reference for distributed computing applications in the field of financial data preprocessing.
References
- [1] Ma, C. Y. (2016). Research on sentiment analysis of online financial information and its relationship with stock market volatility: An implementation based on SparkR (Master's thesis, Hefei University of Technology, Hefei, China).
- [2] Li, Y. (2017). Research on an online big data analysis and decision-making system for smart grids (Master's thesis, North China Electric Power University, Beijing, China).
- [3] Han, J., & Kamber, M. (2012). Data mining: Concepts and techniques (pp. 256–267). Beijing, China: China Machine Press.
- [4] Chen, Z. Y., & Liu, Y. (2021). Anomaly detection and root cause analysis for financial data. Software, 42(7), 119–123.
- [5] Li, Z. (2023). Research on feature selection methods for continuous data and their applications (Master's thesis, China University of Petroleum–Beijing, Beijing, China).
- [6] Xiong, Z. B. (2005). Wavelet denoising preprocessing of financial time series (Master's thesis, Renmin University of China, Beijing, China).
- [7] Tang, Y. (2017). Research on stock linkage effects based on heterogeneous information processing (Master's thesis, Harbin Institute of Technology, Harbin, China).
- [8] Huang, L. M. (2024). Research on massive data analysis methods based on machine learning. China Computer & Communication (Theory Edition), 36(1), 84–86.
- [9] Shi, S. (2021). Research on Spark-based parallel algorithms for finite element clusters (Master's thesis, Hunan University, Changsha, China). https://doi.org/10.27135/d.cnki.ghudu.2021.003799
- [10] National Undergraduate Mathematical Contest in Modeling Organizing Committee. (2023). 2023 Higher Education Press Cup National Undergraduate Mathematical Contest in Modeling, Problem C: Automatic pricing and replenishment decision-making for vegetable commodities [Competition problem].
- [11] Lv, D. H., Tao, Y. Q., Yu, Y. L., et al. (2024). Research on preprocessing methods for satellite on-orbit experimental data. Computer Simulation, 41(5), 62–67.
