Large Language Models and Reinforcement Learning: A Taxonomy of Integration Paradigms, Challenges, and Future Directions
Main Article Content
Keywords
large language models, reinforcement learning, recommend systems, hybrid intelligent systems, sequential decision making
Abstract
This paper examines the integration of large language models (LLMs) and reinforcement learning (RL) in recommender systems, focusing on their theoretical foundations and structural challenges. It highlights the transition from static prediction to sequential decision-making, emphasizing RL’s strengths in long-term reward optimization and interaction modeling, and LLMs’ advantages in semantic understanding and reasoning. Their complementary limitations—RL’s weak semantic representation and LLMs’ lack of long-term optimization—justify their integration. Existing research is classified into “LLM-enhanced RL” and “RL-shaped LLM,” with roles including representation enhancement, reward modeling, policy generation, and environment simulation, under varying coupling levels. The paper proposes a unified three-dimensional framework based on information sources, optimization time scale, and coupling strength, showing that performance differences arise from structural positioning rather than model scale. Key challenges include balancing expressiveness and efficiency, long-term optimization and training stability, and generalization versus specialization. The paper also identifies limitations in evaluation protocols and experimental design, calling for standardized frameworks for long-term value assessment. Overall, integrating LLMs and RL is crucial for advancing recommender systems toward intelligent decision-making agents, with future work focusing on stable coupling and unified evaluation.
References
- [1] H. Ko, S. Lee, Y. Park, and A. Choi, “A survey of recommendation systems: Recommendation models, techniques, and application fields,” Electronics,2022, doi: 10.3390/electronics11010141.
- [2] Lin, Y., Liu, Y., Lin, F., Zou, L., Wu, P., Zeng, W., Chen, H., Miao, C.: A survey on reinforcement learning for recommender systems. IEEE Trans. Neural Netw. Learn. Syst. 34(5), 13164-13184(2023)
- [3] Rezaei, M., Tabrizi, N.: Recommender system using reinforcement learning: a survey. In: Proc. IEEE Int. Conf. Electrical Engineering, Big Data and Algorithms (EEBDA), pp. 148-159 (2022)
- [4] X. Xin, T. Pimentel, A. Karatzoglou, P. Ren, K. Christakopoulou, and Z. Ren, “Rethinking reinforcement learning for recommendation: A prompt perspective,” in Proc. 45th Int. ACM SIGIR Conf. Research and Development in Information Retrieval (SIGIR '22), Madrid, Spain, 2022, pp. 1347-1357.
- [5] M. M. Afsar, T. Crump, and B. Far, “Reinforcement learning based recommender systems: A survey,” ACM Comput. Surv., vol. 55, no. 7, pp. 1-38, Dec. 2022, Art. no. 145.
- [6] K. Bao, J. Zhang, Y. Zhang, W. Wang, F. Feng, and X. He, “TALLRec: An effective and efficient tuning framework to align large language model with recommendation,” in Proc. 17th ACM Conf. Recommender Systems (RecSys '23), Singapore, 2023, pp. 1007-1014.
- [7] B. Sguerra, E. V. Epure, H. Lee, and M. Moussallam, “Biases in LLM-generated musical taste profiles for recommendation,” in Proc. 19th ACM Conf. Recommender Systems (RecSys '25), Prague, Czech Republic, 2025, pp. 527-532.
- [8] Y. Cao et al., “Survey on Large Language Model-Enhanced Reinforcement Learning: Concept, Taxonomy, and Methods,” in IEEE Transactions on Neural Networks and Learning Systems, vol. 36, no. 6, pp. 9737-9757, June 2025, doi: 10.1109/TNNLS.2024.3497992.
- [9] Z. Xue, Q. Cai, B. Yang, L. Hu, P. Jiang, K. Gai, and B. An, “AURO: Reinforcement learning for adaptive user retention optimization in recommender systems,” in Proc. ACM Web Conf. (WWW '25), Sydney, NSW, Australia, 2025, pp. 391-401.
- [10] D. C. Rajapakse and D. Jannach, “Reassessing the effectiveness of reinforcement learning based recommender systems for sequential recommendation,” in Proc. 48th Int. ACM SIGIR Conf. Research and Development in Information Retrieval (SIGIR '25), Padua, Italy, 2025, pp. 3306-3312.
- [11] Z. Ren, N. Huang, Y. Wang, P. Ren, J. Ma, J. Lei, X. Shi, H. Luo, J. Jose, and X. Xin, “Contrastive state augmentations for reinforcement learning-based recommender systems,” in Proc. 46th Int. ACM SIGIR Conf. Research and Development in Information Retrieval (SIGIR '23), Taipei, Taiwan, 2023, pp. 922-931.
- [12] N. Lee and J. Kim, “SEALR: Sequential emotion-aware LLM-based personalized recommendation system,” in Proc. 48th Int. ACM SIGIR Conf. Research and Development in Information Retrieval (SIGIR '25), Padua, Italy, 2025, pp. 2906-2910.
- [13] S. Wang, X. Chen, and L. Yao, “Policy-guided causal state representation for offline reinforcement learning recommendation,” in Proc. ACM Web Conf. (WWW '25), Sydney, NSW, Australia, 2025, pp. 402-412.
- [14] W. Shu, Y. Zeng, Y. Tang, T. Sha, N. Luo, Y. Cheng, X. Liu, F. Zhou, and P. Jiang, “Reward balancing revisited: Enhancing offline reinforcement learning for recommender systems,” in Companion Proc. ACM Web Conf. (WWW Companion '25), Sydney, NSW, Australia, 2025, pp. 1308-1311.
- [15] Y. Zhang, R. Qiu, X. Xu, J. Liu, and S. Wang, “DARLR: Dual-agent offline reinforcement learning for recommender systems with dynamic reward,” in Proc. 48th Int. ACM SIGIR Conf. Research and Development in Information Retrieval (SIGIR '25), Padua, Italy, 2025, pp. 2192-2202.
- [16] Y. Deldjoo, Z. He, J. McAuley, A. Korikov, S. Sanner, A. Ramisa, R. Vidal, M. Sathiamoorthy, A. Kasirzadeh, and S. Milano. A review of modern recommender systems using generative models (gen-recsys). In Proc. 30th ACM SIGKDD Conf. Knowledge Discovery and Data Mining (KDD '24), Barcelona, Spain, 2024, pp. 6448-6458.
- [17] L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems (NeurIPS), 2022, pp. 1-15.
- [18] S. Yun and Y.-k. Lim. User experience with llm-powered conversational recommendation systems: A case of music recommendation. In Extended Abstracts of the CHI Conf. Human Factors in Computing Systems (CHI EA '25), Yokohama, Japan, 2025.
- [19] J. Lin, T. Wang, and K. Qian. Rec-r1: Bridging generative large language models and user-centric recommendation systems via reinforcement learning. arXiv preprint arXiv:2503.24289v4, 2025.
- [20] C. Meng, H. Ling, J. Wang, Y. Liu, S. Zhang, D. Hong, M. Gao, O. Dalal, E. Chi, L. Hong, H. Lu, and N. Han. Balancing fine-tuning and rag: A hybrid strategy for dynamic llm recommendation updates. In Proc. 19th ACM Conf. Recommender Systems (RecSys '25), Prague, Czech Republic, 2025, pp. 919-922.
- [21] N. Ishii and K. Higuchi. A framework for personalized recommendation based on llms with constrained combinatorial optimization. In Extended Abstracts of the CHI Conf. Human Factors in Computing Systems (CHI EA '25), Yokohama, Japan, 2025.
- [22] J. Wang, A. Karatzoglou, I. Arapakis, and J. M. Jose. Reinforcement learning-based recommender systems with large language models for state reward and action modeling. In Proc. 47th Int. ACM SIGIR Conf. Research and Development in Information Retrieval (SIGIR '24), Washington, DC, USA, 2024, pp. 375-385.
- [23] A. Bodaghi, B. C. M. Fung, and K. A. Schmitt. Augmentoxic: Leveraging reinforcement learning to optimize llm instruction fine-tuning for data augmentation to enhance toxicity detection. ACM Transactions on the Web, 19(4), Article 38, 2025.
- [24] F. Liu, H. Guo, X. Li, R. Tang, Y. Ye, and X. He. End-to-End Deep Reinforcement Learning based Recommendation with Supervised Embedding. In Proceedings of the 13th International Conference on Web Search and Data Mining (WSDM ’20), Houston, TX, USA, 2020, pp. 384–392.
- [25] Romain Deffayet, Thibaut Thonet, Jean-Michel Renders, and Maarten de Rijke. 2023. Offline Evaluation for Reinforcement Learning-Based Recommendation: A Critical Issue and Some Alternatives. SIGIR Forum 56, 2 (2023). doi:10.1145/3582900.3582905
- [26] A. Sharma, H. Li, X. Li, and J. Jiao, “Optimizing novelty of top-k recommendations using large language models and reinforcement learning,” in Proc. 30th ACM SIGKDD Conf. Knowledge Discovery and Data Mining (KDD '24), Barcelona, Spain, 2024, pp. 5669–5679.
- [27] X. Zhao, Y. Chen, Z. Wang, H. Zhou, and J. Liu, “Search-R1: Training LLMs to reason and leverage search engines with reinforcement learning,” arXiv preprint arXiv:2503.09516v5, 2025.
