The Combination of RAG and RL: A Technological Breakthrough and Future Directions in Mathematical Reasoning with LLMs
Main Article Content
Keywords
mathematical reasoning, retrieval-augmented generation, reinforcement learning, large language models (LLM), fusion framework
Abstract
Large language models (LLMs) face challenges in mathematical reasoning tasks, including logical disjunction, process hallucination, and a loss of creativity, which severely restricts the reliability and application of LLMs for this purpose. To tackle these issues, this paper applies secondary research, systematically explores the mathematical logic of fusing retrieval-augmented generation and reinforcement learning, and constructs a three-dimensional fusion framework of ‘fusion object –optimization granulari ty–application scenario’. We identify three core technical paths: RL-optimized RAG retrieval strategies, RL-enhanced RAG reasoning chains, and RL-adapted RAG knowledge fusion. According to the results, the precision, sturdiness, and comprehension of mathem atical reasoning can be improved remarkably with the help of this ‘Knowledge Supply-Strategy Optimization’ closed-loop framework. It gets 10% -20% accuracy improvement on the benchmark datasets like MATH and AIME that also controls process hallucination bel ow 5%. This paper further analyzes the limitation of current technologies in terms of cross-domain theorem association, multimodal adaptations, and low-resource scenarios. The article also proposes future research ideas, including agent-based fusion framework and lightweight process supervision, which can efficiently and reliably support mathematics education, scientific research assistance, and engineering computation.
References
- [1] Wang, Z., Zhao, Z. L., & Dou, Z. C.* (2026). ProRAG: Process-Supervised Reinforcement Learning for Retrieval-Augmented Generation. arXiv preprint arXiv:2601.21912v1 [cs.AI], 29 Jan 2026. https://arxiv.org/abs/2601.21912v1
- [2] Forootani, A. (2025). A Survey on Mathematical Reasoning and Optimization with Large Language Models. IEEE Transactions on Artificial Intelligence, Vol. 00, No. 0. arXiv:2503.17726v1 [cs.AI].
- [3] Hubert, T., Mehta, R., Sartran, L., et al. (2026). Olympiad-level formal mathematical reasoning with reinforcement learning. Nature, 651, 607-613. https://doi.org/10.1038/s41586-025-09833-y
- [4] Guo C, Huang J J, Xie H J, et al. LiR³ AG: A Lightweight Rerank Reasoning Strategy Framework for Retrieval-Augmented Generation [EB/OL]. (2026-01-19) https://arxiv.org/abs/2512.18329v2.
- [5] Lan, Y., Xu, Y., & Chen, H. (2026). Template-Theorems Graph Construction to Enhance Mathematical Reasoning Capabilities of LLM. Proceedings of the 40th AAAI Conference on Artificial Intelligence (AAAI-26)
- [6] Gupta, S., Ranjan, R., & Singh, S. N. (2024). A Comprehensive Survey of Retrieval-Augmented Generation (RAG): Evolution, Current Landscape and Future Directions. arXiv preprint.
- [7] DeepSeek Team. (2026). DeepSeek-R1 Practice: How to train a math-superior AI model with pure reinforcement learning. CSDN Blog. https://www.csdn.net/.
- [8] Arabzadeh, N., Ma, W., Min, S., & Zaharia, M. (2025). Restructuring the corpus makes RAG work for math. 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: MATH-AI.
- [9] Dixit, P., & Oates, T. (2024). SBI-RAG: Enhancing math word problem solving for students through schema-based instruction and retrieval-augmented generation. 38th Conference on Neural Information Processing Systems (NeurIPS 2024).
- [10] Li, Y., Zhang, W., Yang, Y., et al. (2025). Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs. arXiv preprint arXiv:2507.09477v2 [cs.CL].
