Evaluating AI “Understanding” with Cognitive-Psychological Criteria: Evidence, Gaps, and Structural Limits
Main Article Content
Keywords
cognitive psychology, artificial intelligence, understanding ability
Abstract
With the rapid development of large-scale language models, the performance of artificial intelligence in language understanding and reasoning tasks has become increasingly strong. As a result, many people have gradually regarded the success of tasks as the same as true understanding. However, from the perspective of cognitive psychology, understanding is not merely about giving correct or fluent responses; it is actually an internal psychological process. It includes aspects such as meaning construction, context model update, reasoning generation, and metacognitive regulation. Under such circumstances, this paper systematically studies the cognitive understanding ability of artificial intelligence from the perspective of cognitive psychology. It employs the methods of literature review and qualitative case analysis. First, it clarifies the concept definitions and evaluation criteria of understanding in cognitive psychology by combining classic theories and experimental evidence, and then uses these criteria as the analytical framework. Study the performance of “similar understanding” in contemporary artificial intelligence systems. It was found that although artificial intelligence can approach human understanding at the behavioral level, it has not reached the core cognitive psychological standards. AI lacks experience-based situational models, causal and goal-oriented reasoning mechanisms, and inherent metacognitive monitoring. These limitations are structural, not quantitative issues. It reflects the fundamental difference between artificial intelligence systems and human cognition. This article can help clarify the theoretical distinction between surface manifestations and true understanding, and also provide a cognitive psychology foundation for a more cautious explanation of artificial intelligence capabilities.
References
- [1] Russell, S. J., Norvig, P. Artificial intelligence: A modern approach. 3rd ed. Upper Saddle River, NJ: Prentice Hall; 2010.
- [2] LeCun, Y., Bengio, Y., Hinton, G. Deep learning. Nature. 2015, 521(7553), 436-444. https://doi.org/10.1038/nature14539
- [3] Stanford Institute for Human-Centered Artificial Intelligence. The AI Index 2024 Annual Report. Stanford University; 2024. Available from: https://aiindex.stanford.edu/report/ (accessed 7 June 2026).
- [4] Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., et al. Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems 33. Curran Associates; 2020, pp. 1877-1901. https://doi.org/10.48550/arXiv.2005.14165
- [5] China Academy of Information and Communications Technology. Research Report on the Development of Artificial Intelligence Industry (2025). Beijing: China Academy of Information and Communications Technology; 2026. (in Chinese)
- [6] Anderson, J. R. Cognitive psychology and its implications. 8th ed. New York, NY: Worth Publishers; 2014.
- [7] Gardner, H. The mind's new science: A history of the cognitive revolution. New York, NY: Basic Books; 1985.
- [8] Searle, J. R. Minds, brains, and programs. Behavioral and Brain Sciences. 1980, 3(3), 417-457. https://doi.org/10.1017/S0140525X00005756
- [9] Harnad, S. The symbol grounding problem. Physica D: Nonlinear Phenomena. 1990, 42(1-3), 335-346. https://doi.org/10.1016/0167-2789(90)90087-6
- [10] Levesque, H. J., Davis, E., Morgenstern, L. The Winograd Schema Challenge. In Proceedings of the Thirteenth International Conference on Principles of Knowledge Representation and Reasoning. AAAI Press; 2012, pp. 552-561.
- [11] Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., Bowman, S. R. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP. Association for Computational Linguistics; 2018, pp. 353-355. https://doi.org/10.18653/v1/W18-5446
- [12] Bender, E. M., Koller, A. Climbing towards NLU: On meaning, form, and understanding in the age of data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics; 2020, pp. 5185-5198. https://doi.org/10.18653/v1/2020.acl-main.463
- [13] Bender, E. M., Gebru, T., McMillan-Major, A., Shmitchell, S. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency. ACM; 2021, pp. 610-623. https://doi.org/10.1145/3442188.3445922
- [14] Kintsch, W. Comprehension: A paradigm for cognition. Cambridge: Cambridge University Press; 1998.
- [15] Zwaan, R. A. From Words to Worlds: Twenty-Five Years of Advances in Situation Model Research. Current Directions in Psychological Science. 2025, 34(5), 287-292. https://doi.org/10.1177/09637214251326812
- [16] Zwaan, R. A., Radvansky, G. A. Situation models in language comprehension and memory. Psychological Bulletin. 1998, 123(2), 162-185. https://doi.org/10.1037/0033-2909.123.2.162
- [17] Zwaan, R. A., Magliano, J. P., Graesser, A. C. Dimensions of situation model construction in narrative comprehension. Journal of Experimental Psychology: Learning, Memory, and Cognition. 1995, 21(2), 386-397. https://doi.org/10.1037/0278-7393.21.2.386
- [18] McNamara, D. S., Magliano, J. P. Toward a comprehensive model of comprehension. In The psychology of learning and motivation. Elsevier; 2009, Vol. 51, pp. 297-384. https://doi.org/10.1016/S0079-7421(09)51009-2
- [19] Tibken, C., Richter, T., Wannagat, W. Metacognitive comprehension monitoring: Cognitive abilities explain performance differences between younger and older adults. Scientific Studies of Reading. 2024, 28(3), 284-302. https://doi.org/10.1080/10888438.2023.2261572
- [20] Wannagat, W., Nieding, G., Tibken, C. Age-related decline of metacognitive comprehension monitoring in adults aged 50 and older: Effects of cognitive abilities and educational attainment. Cognitive Development. 2024, 70, Article 101440. https://doi.org/10.1016/j.cogdev.2024.101440
- [21] Keller, T. A., Mason, R. A., Legg, A. E., Just, M. A. The neural and cognitive basis of expository text comprehension. npj Science of Learning. 2024, 9, Article 21. https://doi.org/10.1038/s41539-024-00232-y
- [22] Graesser, A. C., Singer, M., Trabasso, T. Constructing inferences during narrative text comprehension. Psychological Review. 1994, 101(3), 371-395. https://doi.org/10.1037/0033-295X.101.3.371
- [23] Sun, C. Y., Peng, P., Chen, H. X., et al. The Effects of the Processing and the Storage in English and Chinese Working Memory on English Reading cognitive Comprehension among Chinese Middle Sehool Students. Psychological Development and Education. 2012, 28(1), 61-69. https://doi.org/10.16187/j.cnki.issn1001-4918.2012.01.003 (in Chinese)
- [24] Liu, W. F., Si, J. W., Wang, Y. X. Metacognitive factors in cognitive strategy selection. Advances in Psychological Science. 2011, (7), 10-14. (in Chinese)
- [25] Dunlosky, J., Lipko, A. R. Metacomprehension: A brief history and how to improve its accuracy. Current Directions in Psychological Science. 2007, 16(4), 228-232. https://doi.org/10.1111/j.1467-8721.2007.00509.x
- [26] Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems. 2022, 35, 27730-27744. https://doi.org/10.48550/arXiv.2203.02155
