A Survey: Lossless Model Accuracy Data Partitioning and Training Parallelism Strategies for Distributed Machine Learning
Main Article Content
Keywords
distributed machine learning, accuracy losslessness, non-independent and identically distributed, gradient compensation
Abstract
In the digital era of big data and large-scale models, conventional stand-alone computing power and storage approach fails to meet the demands of large-scale machine learning models, therefore, distributed data parallelism has been the key supporting technology for breaking hardware bottleneck. However, current distributed training is facing numerous actual conflicts in the practical applications: data partitioning and resource scheduling, heterogeneous clusters and global synchronization. This passage elaborates on the practical dilemmas of distributed machine learning training, achieving thorough integration of data partitioning patterns, parallel operation architectures and precision stabilization strategies by three dimensions: problems, methods and findings. The study points out that existing technology manifests multiple limitations, such as inefficient training iteration, poor flexibility of static partitioning, and the difficulty in simultaneously maintaining accuracy invariance and convergence consistency. Subsequent research can develop with the direction of building Non-Independent and Identically Distributed (Non-IID) adaptive partitioning mechanisms, constructing efficient resource coordination platforms and correcting local gradients, which provides reliable references for establishing high-precision distributed machine learning systems.
References
- [1] Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., & Catanzaro, B.(2019). Megatron-LM: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv: 1909. 08053.
- [2] Dean, J., & Ghemawat, S.(2008). MapReduce: Simplified data processing on large clusters. Communications of the ACM, 51(1), 107–113.
- [3] Um, T., et al.(2024). Metis: Fast Automatic Distributed Training on Heterogeneous GPUs. 2024 USENIX Annual Technical Conference(USENIX ATC 24).
- [4] Dean, J., & Barroso, L. A.(2013). The tail at scale. Communications of the ACM, 56(2), 74-80.
- [5] Karimireddy, S. P., Kale, S., Mohri, M., Reddi, S., Stich, S., & Suresh, A. T.(2020). Scaffold: Stochastic controlled averaging for federated learning. International Conference on Machine Learning(ICML), 5132- 5143.
- [6] Um, T., et al.(2024). Metis: Fast Automatic Distributed Training on Heterogeneous GPUs. USENIX Annual Technical Conference(ATC).
- [7] Alistarh, D., Grubic, D., Li, J., Tomioka, R., & Vojnovic, M.(2017). QSGD: Communication-efficient SGD via gradient quantization and encoding. Advances in Neural Information Processing Systems(NeurIPS).
- [8] Qiao, A., Choe, S. K., Subramanya, S. J., Neiswanger, W., Peng, Q., Xing, E., &Ganger, G. R.(2021). Pollux: Co-adaptive cluster scheduling for goodput-optimized deep learning. In 15th USENIX Symposium on Operating Systems Design and Implementation(OSDI 21)(pp. 1-18).
- [9] Peng, Y., Zhu, Y., Chen, Y., Bao, Y., Yi, B., Lan, C.,...&Guo, C.( 2019). A generic communication scheduler for distributed DNN training acceleration. In Proceedings of the 27th ACM Symposium on Operating Systems Principles(SOSP '19)(pp. 16-29).
- [10] Li, T., Sahu, A. K., Zaheer, M., Sanjabi, M., Talwalkar, A., & Smith, V.(2020). Federated optimization in heterogeneous networks. Proceedings of Machine Learning and Systems(MLSys), 2, 429-450.
- [11] Chai, Z., Ali, A., Zawad, S., Truex, S., Anwar, A., Baracaldo, N.,...&Yan, F.(2020). TiFL: A tier-based federated learning system. In Proceedings of the 29th International Symposium on High-Performance Parallel and Distributed Computing(HPDC)(pp. 125-136).
- [12] Stich, S. U., Cordonnier, J. B., &Jaggi, M.(2018). Sparsified SGD with memory. Advances in Neural Information Processing Systems(NeurIPS), 31.
