Progress and Analysis of Optimization for Large Language Models Based on Reinforcement Learning
Main Article Content
Keywords
large language model, reinforcement learning, reasoning ability
Abstract
Large language models (LLMs) are one of the current research focuses in society and have extensive applications in various fields. However, when facing complex tasks such as mathematical reasoning at present, it will exhibit problems such as weak generalization ability. Reinforcement Learning (RL) can effectively optimize these issues through reward-guided strategies, thus becoming the core technical path for enhancing the reasoning capabilities of large language models. This article systematically reviews and analyzes the boundaries of how reinforcement learning optimizes the reasoning capabilities of large language models under different resource and scale constraints, as well as the adaptability of reinforcement learning's optimization of large language models in different scenarios. The analysis shows that reinforcement lea rning algorithms are beneficial for aligning model reasoning and improving the reasoning chain of the model; based on reinforcement learning, strategies such as data minimization training and small model optimization can enhance the resource utilization ef ficiency of the model during training; however, the improvement of reinforcement learning algorithms on large language models still has problems such as difficulties in expanding the reasoning boundaries.
References
- [1] Singh, J., Magazine, R., Pandya, Y., & Nambi, A. (2025). Agentic reasoning and tool integration for llms via reinforcement learning. arXiv preprint arXiv:2505.01441.
- [2] Narvekar, S., & Stone, P. (2018). Learning curriculum policies for reinforcement learning. arXiv preprint arXiv:1812.00285.
- [3] Guo, D., Yang, D., Zhang, H., Song, J., Wang, P., Zhu, Q.,... & He, Y. (2025). Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948.
- [4] Wu, J., Ning, L., Liu, L., Lee, H., Wu, N., Wang, C., Prakash, S., O’Banion, S., Green, B., & Xie, J. (2025). RLPF: Reinforcement Learning from Prediction Feedback for User Summarization with LLMs. Proceedings of the AAAI Conference on Artificial Intelligence, 39(24), 25488-25496.
- [5] Dang, Q. A., & Ngo, C. (2025). Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't. arXiv preprint arXiv:2503.16219.
- [6] Wang, Y., Yang, Q., Zeng, Z., Ren, L., Liu, L., Peng, B.,... & Shen, Y. (2025). Reinforcement learning for reasoning in large language models with one training example. arXiv preprint arXiv:2504.20571.
- [7] Tan, Z., Geng, H., Yu, X., Zhang, M., Wan, G., Zhou, Y.,... & Bai, L. (2025). Scaling behaviors of llm reinforcement learning post-training: An empirical study in mathematical reasoning. arXiv preprint arXiv:2509.25300.
- [8] Pan, P. C., Liang, Y., & Lin, S. (2026). Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation. arXiv preprint arXiv:2602.09305.
- [9] Chen, M., Sun, L., Li, T., Sun, H., Zhou, Y., Zhu, C.,... & Chen, W. (2025). Learning to reason with search for llms via reinforcement learning. arXiv preprint arXiv:2503.19470.
- [10] Lan, G., Inan, H. A., Abdelnabi, S., Kulkarni, J., Wutschitz, L., Shokri, R.,... & Sim, R. (2025). Contextual integrity in LLMs via reasoning and reinforcement learning. arXiv preprint arXiv:2506.04245.
- [11] Yue, Y., Chen, Z., Lu, R., Zhao, A., Wang, Z., Song, S., & Huang, G. (2025). Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?. arXiv preprint arXiv:2504.13837.
- [12] Cheng, Z., Hao, S., Liu, T., Zhou, F., Xie, Y., Yao, F.,... & Hu, Z. (2025). Revisiting reinforcement learning for llm reasoning from a cross-domain perspective. arXiv preprint arXiv:2506.14965.
