Reasoning in Large Language Models: Methods, Challenges, and Future Directions

Main Article Content

Zhuangchen Xu

Keywords

large language models, reasoning, chain-of-thought, tool-augmented models, AI agents

Abstract

In recent years, large language models have developed rapidly and are now widely used in many everyday tasks, such as conversation, information retrieval, and code generation. Although these models can produce fluent and coherent text, their ability to perform reliable reasoning is still limited, especially in tasks that require multiple steps or logical consistency. This paper reviews several methods aimed at improving reasoning in LLMs. It first introduces prompt-based approaches, including Chain-of-Thought, self-consistency, and Auto- CoT, which try to guide models to generate intermediate reasoning steps. It then discusses search-based methods, such as Tree-of-Thought, in which reasoning is treated as a process of exploring multiple paths. In addition, the paper examines the hallucination problem and the use of retrieval-based techniques to improve factual accuracy. Benchmark datasets for evaluating reasoning ability are also briefly discussed. Furthermore, the paper examines recent developments in agent-based reasoning and tool use, in wh ich models can interact with external systems and perform more complex tasks. While these methods show some improvements, several open challenges remain, including issues of reasoning reliability, error accumulation, and evaluation.

Abstract 8 | PDF Downloads 3

References

  • [1] Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E.,... & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35, 24824-24837.
  • [2] Zhang, Z., Zhang, A., Li, M., & Smola, A. (2022). Automatic chain of thought prompting in large language models. arXiv preprint arXiv:2210.03493.
  • [3] Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., & Narasimhan, K. (2023). Tree of thoughts: Deliberate problem solving with large language models. Advances in neural information processing systems, 36, 11809-11822.
  • [4] Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S.,... & Zhou, D. (2022). Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
  • [5] Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., & Iwasawa, Y. (2022). Large language models are zero-shot reasoners. Advances in neural information processing systems, 35, 22199-22213.
  • [6] Chen, J., Chen, L., Huang, H., & Zhou, T. (2023). When do you need Chain-of-Thought Prompting for ChatGPT?. arXiv preprint arXiv:2304.03262.
  • [7] Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N.,... & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems, 33, 9459-9474.
  • [8] Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P.,... & Lowe, R. (2022). Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35, 27730-27744.
  • [9] Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P.,... & Amodei, D. (2020). Language models are few-shot learners. Advances in neural information processing systems, 33, 1877 - 1901.
  • [10] Ferrag, M. A., Tihanyi, N., & Debbah, M. (2025). From llm reasoning to autonomous ai agents: A comprehensive review. arXiv preprint arXiv:2504.19678.
  • [11] Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., & Steinhardt, J. (2020). Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300.
  • [12] Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L.,... & Schulman, J. (2021). Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  • [13] Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E.,... & Sharick, E. (2023). Paper Review:'Sparks of Artificial General Intelligence: Early experiments with GPT-4'.
  • [14] Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Hambro, E.,... & Scialom, T. (2023). Toolformer: Language models can teach themselves to use tools. Advances in neural information processing systems, 36, 68539-68551.
  • [15] Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K. R., & Cao, Y. (2022, October). React: Synergizing reasoning and acting in language models. In The eleventh international conference on learning representations.
  • [16] Qin, Y., Hu, S., Lin, Y., Chen, W., Ding, N., Cui, G.,... & Sun, M. (2024). Tool learning with foundation models. ACM Computing Surveys, 57(4), 1-40.
  • [17] Xi, Z., Chen, W., Guo, X., He, W., Ding, Y., Hong, B.,... & Gui, T. (2025). The rise and potential of large language model based agents: A survey. Science China Information Sciences, 68(2), 121101.