Factual Hallucination in Medical Large Language Models: Typology, Evaluation, and Systematic Governance

Main Article Content

Zijiao Liu

Keywords

large language models, medical hallucination, retrieval-augmented generation, patient safety, clinical governance, risk stratification

Abstract

Large language models (LLMs) have demonstrated transformative potential in clinical documentation generation, diagnostic assistance, and patient consultation. However, their tendency toward “hallucination”— generating semantically fluent but factually inconsistent content with established medical knowledge or input context—constitutes a core safety barrier to clinical deployment. This paper systematically reviews the typology, evaluation frameworks, underlying mechanisms, and mitigation strategies for medica l LLM hallucinations. First, we propose a fine-grained hallucination classification framework based on medical knowledge graphs and establish a three-tier risk stratification scheme that references FDA medical AI software risk classification standards. Sec ond, we analyze multi-factor causes from data, model, and inference dimensions, highlighting the unique characteristics of the medical domain. Third, we systematically review current evaluation benchmarks and methodologies, deeply analyzing the limitations of automated metrics and the potential of LLMs as evaluators. Regarding mitigation strategies, we propose a three-stage classification framework—”training-stage internal intervention —inference-stage external constraints —post-processing collaborative verification” —that provides in-depth analysis of each strategy's internal mechanisms and trade- offs. Finally, we discuss challenges in real-world clinical deployment, emphasizing that the “human-in-the- loop” paradigm remains an irreplaceable safety line.

Abstract 17 | PDF Downloads 5

References

  • [1] Jung, K. H. (2025). Large language models in medicine: clinical applications, technical challenges, and ethical considerations. Healthcare Informatics Research, 31(2), 114-124.
  • [2] Li, S., Wang, X., Chen, Y., Tian, M., Lin, P., Lai, M., & Jiang, L. (2026). Large language models for primary care ophthalmic education: a systematic review. Frontiers in Medicine, 13, 1810098.
  • [3] Zhang, R. (2025). Classification and solution of large language model hallucination. In ITM Web of Conferences(Vol. 80, p. 01035). EDP Sciences.
  • [4] Passban, P., Matthews, A., Roosta, T., & Vankadaru, V. (2026). Rethinking Medical LLM Hallucinations: A System-Level Survey.
  • [5] Asgari, E., Montañ a-Brown, N., Dubois, M., Khalil, S., Balloch, J., Yeung, J. A., & Pimenta, D. (2025). A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. NPJ digital medicine, 8(1), 274.
  • [6] Yoon, S. M., Lyu, J., Djunadi, T. A., Song, J., Kim, H. S., Min, R. S.,... & Chae, Y. K. (2025). Navigating artificial intelligence (AI) accuracy: A meta-analysis of hallucination incidence in large language model (LLM) responses to oncology questions.
  • [7] Prandner, D., Wetzelhü tter, D., & Hese, S. (2025, January). ChatGPT as a data analyst: an exploratory study on AI-supported quantitative data analysis in empirical research. In Frontiers in Education (Vol. 9, p. 1417900). Frontiers Media SA.
  • [8] Lu, J., Liu, J., Zheng, X., Yang, M., Wang, J., Wang, P., & Zhang, Y. (2026, March). MHB: Medical Hallucination Benchmark for Large Language Models in Complex Clinical Tasks. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 40, No. 45, pp. 38971-38978).
  • [9] Ross, J. (2019). Be Aware of Health Technology Hazards. Journal of PeriAnesthesia Nursing, 34(2), 435-438.
  • [10] Li, Y., Fu, X., Verma, G., Buitelaar, P., & Liu, M. (2025). Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems. arXiv preprint arXiv:2510.24476.
  • [11] Kuculmez, O., Usen, A., & Ahi, E. D. (2025). Referential hallucination and clinical reliability in large language models: a comparative analysis using regenerative medicine guidelines for chronic pain. Rheumatology International, 45(10), 240.
  • [12] Salehi, S., Singh, Y., Horst, K. K., Hathaway, Q. A., & Erickson, B. J. (2025). Agentic AI and Large Language Models in Radiology: Opportunities and Hallucination Challenges. Bioengineering, 12(12), 1303.
  • [13] Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H.,... & Liu, T. (2025). A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, 43(2), 1-55.
  • [14] Aydin, S., Karabacak, M., Vlachos, V., & Margetis, K. (2024). Large language models in patient education: a scoping review of applications in medicine. Frontiers in medicine, 11, 1477898.
  • [15] Pandit, S., Xu, J., Hong, J., Wang, Z., Chen, T., Xu, K., & Ding, Y. (2025, November). Medhallu: A comprehensive benchmark for detecting medical hallucinations in large language models. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (pp. 2858-2873).
  • [16] Gu, Z., Chen, J., Liu, F., Yin, C., & Zhang, P. (2026). MedVH: Toward systematic evaluation of hallucination for large vision language models in the medical context. Advanced Intelligent Systems, 8(1), 2500255.
  • [17] Li, H., Huang, J., Ji, M., Yang, Y., & An, R. (2025). Use of retrieval-augmented large language model for COVID-19 fact-checking: development and usability study. Journal of medical Internet research, 27, e66098.
  • [18] Lee, J. T., Li, V. C. S., Wu, J. J., Chen, H. H., Su, S. S. Y., Chang, B. P. H.,... & Atun, R. (2025). Evaluation of performance of generative large language models for stroke care. npj Digital Medicine, 8(1), 481.
  • [19] Garcia-Fernandez, C., Felipe, L., Shotande, M., Zitu, M., Tripathi, A., Rasool, G.,... & Valdes, G. (2025). Trustworthy ai for medicine: Continuous hallucination detection and elimination with check. arXiv preprint arXiv:2506.11129.
  • [20] Wang, Y., Gao, S., Liu, J., Jiang, S., Haoxiang, X., Zhang, X.,... & Liu, Z. (2026, March). Beyond n-grams: A hierarchical reward learning framework for clinically-aware medical report generation. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 40, No. 40, pp. 33719-33727).
  • [21] Deng, Y., Zhao, S., Miao, Y., Zhu, J., & Li, J. (2025). MedKA: A knowledge graph-augmented approach to improve factuality in medical Large Language Models. Journal of Biomedical Informatics, 104871.
  • [22] Akl, A., Khamis, A., Cheraghian, A., Wang, Z., Khalifa, S., & Wang, K. (2026). HIME: Mitigating Object Hallucinations in LVLMs via Hallucination Insensitivity Model Editing. arXiv preprint arXiv:2602.18711.
  • [23] Xu, S., Yan, Z., Dai, C., & Wu, F. (2025). MEGA-RAG: a retrieval-augmented generation framework with multi-evidence guided answer refinement for mitigating hallucinations of LLMs in public health. Frontiers in Public Health, 13, 1635381.
  • [24] Scanlin, J., Cuesta, A., & Varsavsky, M. (2026). Representation Before Retrieval: Structured Patient Artifacts Reduce Hallucination in Clinical AI Systems. medRxiv, 2026-02.
  • [25] Wu, J., Wu, X., Zheng, Y., & Yang, J. (2025). Clinical pathway-aware large language models for reliable and transparent medical dialogue. Journal of Biomedical Informatics, 104942.
  • [26] Yang, D., Wei, J., Li, M., Liu, J., Liu, L., Hu, M.,... & Zhang, L. (2025). MedAide: information fusion and anatomy of medical intents via LLM-based agent collaboration. Information Fusion, 103743.
  • [27] Garcia-Fernandez, C., Felipe, L., Shotande, M., Zitu, M., Tripathi, A., Rasool, G.,... & Valdes, G. (2025). Trustworthy ai for medicine: Continuous hallucination detection and elimination with check. arXiv preprint arXiv:2506.11129.