Mitigating Hallucinations in Natural Language Generation through Prompt Engineering: A Mechanism- Oriented Narrative Review
Main Article Content
Keywords
large language models, hallucination, prompt engineering, chain-of-verification, retrieval-augmented generation, information retrieval
Abstract
Large language models (LLMs) can generate fluent and confident responses that are factually incorrect, unsupported by evidence, or inconsistent with the source material. These hallucinations reduce the reliability of LLM-based question answering, summarization, dialogue, and information retrieval systems, especially when users rely on generated content for consequential decisions. This narrative review examines prompt engineering as a practical inference-time approach for reducing hallucinations without modifying model parameters. Instead of listing individual prompting templates, the paper organizes existing methods according to five functional mechanisms: constraint and evidence grounding, decomposition and verification, multi-path consistency, iterative refinement and tool use, and retrieval-augmented prompting. For each mechanism, the review discusses the hallucination types it is most likely to address, the assumptions required for success, and the conditions under which it may fail. Particular attention is given to Chain-of-Verification (CoVE), which improves auditability by separating answer generation from targeted checking, but remains vulnerable when verification is performed by the same model without independent evidence. The review argues that prompt engineering should be understood as process-level risk reduction rather than a complete solution. Its strongest use is within evidence-centered system designs that combine retrieval, claim-level verification, calibrated abstention, provenance display, and human review in high-risk contexts.
References
- [1] Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55 (12), 1–38. https://doi.org/10.1145/3571730
- [2] Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., & Liu, T. (2025). A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, 43 (2), 1–55. https://doi.org/10.1145/3703155
- [3] Zhang, Y., Li, Y., Cui, L., Cai, D., Liu, L., Fu, T., Huang, X., Zhao, E., Zhang, Y., Chen, Y., Wang, L., Luu, A. T., Bi, W., & Shi, S. (2025). Siren's song in the AI ocean: A survey on hallucination in large language models. Computational Linguistics, 51(4), 1373–1418. https://doi.org/10.1162/coli.a.16
- [4] Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., & Neubig, G. (2023). Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Computing Surveys, 55(9), 1–35. https://doi.org/10.1145/3560815
- [5] Sahoo, P., Singh, A. K., Saha, S., Jain, V., Mondal, S., & Chadha, A. (2024). A systematic survey of prompt engineering in large language models: Techniques and applications. arXiv. https://arxiv.org/abs/2402.07927
- [6] Dhuliawala, S., Komeili, M., Xu, J., Raileanu, R., Li, X., Celikyilmaz, A., & Weston, J. (2023). Chain-of-verification reduces hallucination in large language models. arXiv. https://arxiv.org/abs/2309.11495
- [7] Simhi, A., Itzhak, I., Barez, F., Stanovsky, G., & Belinkov, Y. (2025). Trust me, I'm wrong: High-certainty hallucinations in LLMs. arXiv. https://arxiv.org/abs/2502.12964
- [8] Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E. H., Le, Q. V., & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems (Vol. 35, pp. 24824–24837).
- [9] Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E. H., Narang, S., Chowdhery, A., & Zhou, D. (2023). Self-consistency improves chain of thought reasoning in language models. In Proceedings of the 11th International Conference on Learning Representations (ICLR).
- [10] Chia, Y. K., Chen, G., Tuan, L. A., Poria, S., & Bing, L. (2023). Contrastive chain-of-thought prompting. arXiv. https://arxiv.org/abs/2311.09277
- [11] Manakul, P., Liusie, A., & Gales, M. J. F. (2023). SelfCheckGPT: Zero-resource black-box hallucination detection for generative large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 9004–9017). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.emnlp-main.557
- [12] Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., Gupta, S., Majumder, B. P., Hermann, K., Welleck, S., Yazdanbakhsh, A., & Clark, P. (2023). Self-refine: Iterative refinement with self-feedback. arXiv. https://arxiv.org/abs/2303.17651
- [13] Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. In Proceedings of the 11th International Conference on Learning Representations (ICLR).
- [14] Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Kü ttler, H., Lewis, M., Yih, W., Rocktä schel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems (Vol. 33, pp. 9459–9474).
- [15] Yu, W., Zhang, H., Pan, X., Ma, K., Wang, H., & Yu, D. (2024). Chain-of-note: Enhancing robustness in retrieval-augmented language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 14672–14685). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-main.813
- [16] He, B., Chen, N., He, X., Yan, L., Wei, Z., Luo, J., & Ling, Z. -H. (2024). Retrieving, rethinking and revising: The chain-of-verification can improve retrieval augmented generation. In Findings of the Association for Computational Linguistics: EMNLP 2024 (pp. 10371–10393). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-emnlp.607
