Beyond Fluency: Human Verification of Safety-Critical Errors in Generative-AI First-Aid Translation
Main Article Content
Keywords
generative AI translation, first-aid translation, safety-critical errors, fluency-detectability gap, MQM, skopos theory, translator as safety firewall, limited English proficiency, human verification
Abstract
The rapid integration of generative artificial intelligence (GenAI) into healthcare communication has introduced machine translation (MT) into time-critical, high-stakes scenarios where limited English proficient (LEP) individuals must act on translated instructions without professional mediation. While GenAI systems produce remarkably fluent outputs, this fluency can mask errors that fatally undermine the communicative purpose of first-aid texts—the skopos of enabling correct life-saving actions. This study identifies and empirically substantiates a previously unnamed phenomenon: the fluency-detectability gap, a class of safety-critical errors that are highly severe yet poorly detectable because they surface as fluent, internally coherent, and plausible target-language text. These errors systematically evade three lines of defense: (i) fluency-oriented automatic quality estimation metrics, (ii) monolingual end-users lacking MT literacy, and (iii) even clinical reviewers who do not understand the target language. Through a mixed-methods product analysis with three-layer triangulation, we examine English-to-Chinese translations of public first-aid materials (CPR, choking, severe bleeding, burns, anaphylaxis, poisoning, and seizures) generated by GPT-4o, DeepL, and Google Translate. Using an adapted MQM (Multidimensional Quality Metrics) framework augmented with a detectability dimension, we demonstrate that professional translators—anchored to the source text and bilingual in situ—consistently intercept critical-yet-invisible errors that automatic metrics and monolingual users fail to detect. We reconceptualize translator verification as a loyalty-based safety intervention [1] rather than mere post-editing, and propose the severity×detectability matrix as a reusable quality-control instrument for safety-critical translation. Our findings challenge fluency-centered evaluation paradigms and advocate for mandatory bilingual human-in-the-loop verification in GenAI-mediated first-aid communication.
References
- [1] Nord, C. (2018). Translating as a purposeful activity: Functionalist approaches explained (2nd ed.). Routledge.
- [2] Jiao, W., Wang, W., Huang, J., Wang, X., & Tu, Z. (2023). Is ChatGPT a good translator? A preliminary study. arXiv preprint arXiv:2301.08745.
- [3] Wang, L., & Bu, H. (2022). Emergency language services and emergency translation: Current situation and reflections. Foreign Languages in China, 19(4), 76–83. (in Chinese).
- [4] Khoong, E. C., Steinbrook, E., Brown, C., & Fernandez, A. (2019). Assessing the use of Google Translate for Spanish and Chinese translations of emergency department discharge instructions. JAMA Internal Medicine, 179(4), 580–582.
- [5] Kong, M., Fernandez, A., Bains, J., & Khoong, E. C. (2025). Evaluation of the accuracy and safety of machine translation of patient-specific discharge instructions: A comparative study. Journal of General Internal Medicine, 40(2), 312–320.
- [6] American Red Cross. (2023). First aid/CPR/AED participant’s manual. American Red Cross.
- [7] World Health Organization. (2021). Basic emergency care: Approach to the acutely ill and injured. World Health Organization.
- [8] American Heart Association. (2020). CPR and ECC guidelines. American Heart Association.
- [9] St. John Ambulance. (2022). First aid reference guide. St. John Ambulance.
- [10] Reiss, K., & Vermeer, H. J. (2013). Towards a general theory of translational action: Skopos theory explained (C. Nord, Trans.). Routledge. (Original work published 1984)
- [11] Rodriguez, J., Cohen, R., & Lagu, T. (2021). Machine translation in clinical settings: A review of risks and practice patterns. Journal of Patient Safety, 17(5), 402–408.
- [12] Specia, L., Paetzold, G. H., & Scarton, C. (2018). Multi-level translation quality prediction. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (pp. 3155–3165). Association for Computational Linguistics.
- [13] Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., ... & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), 1–38.
- [14] Raunak, V., Kumar, V., & Khadilkar, H. (2021). A study of hallucinations in neural machine translation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Student Research Workshop (pp. 23–33). Association for Computational Linguistics.
- [15] Al Sharou, K., & Specia, L. (2022). A taxonomy and study of critical errors in machine translation. In Proceedings of the 23rd Annual Conference of the European Association for Machine Translation (pp. 171–180). European Association for Machine Translation.
- [16] Grice, H. P. (1975). Logic and conversation. In P. Cole & J. L. Morgan (Eds.), Syntax and semantics, Vol. 3: Speech acts (pp. 41–58). Academic Press.
- [17] Vieira, L. N., O’Hagan, M., & O’Sullivan, C. (2021). Understanding the societal impacts of machine translation: A critical review of the literature on medical and legal use cases. Information, Communication & Society, 24(11), 1515–1532.
- [18] Castilho, S., Moorkens, J., Gaspari, F., Doherty, S., & Way, A. (2018). Is machine translation usable? A review of the literature. Translation Spaces, 7(1), 16–47.
- [19] Bowker, L., & Ciro, J. B. (2019). Machine translation and global research: Towards a machine translation literacy. Emerald Publishing.
- [20] Flores, G., Abreu, M., & Barrueco, S. (2022). Bilingual staff in healthcare: A qualitative study of translation practices. Journal of Immigrant and Minority Health, 24(3), 642–650.
- [21] Reiss, K. (1976). Texttyp und Übersetzungsmethode: Der operative Text. Scriptor.
- [22] Zhang, M. (2005). Function plus loyalty: An introduction to Christiane Nord’s functionalist translation theory. Journal of Foreign Languages, (1), 60–65. (in Chinese).
- [23] Lommel, A., Uszkoreit, H., & Burchardt, A. (2014). Multidimensional Quality Metrics (MQM): A framework for declaring and describing translation quality metrics. Traduimática, 12, 455–463.
- [24] Mariana, V., Cox, T., & Melby, A. (2015). The Multidimensional Quality Metrics (MQM) framework: A new framework for translation quality assessment. The Journal of Specialised Translation, 23, 137–161.
- [25] OpenAI. (2023). GPT-4 technical report. arXiv preprint arXiv:2303.08774.
- [26] Cui, Q. (2021). Translation technology and human–machine coupling in the age of artificial intelligence. Chinese Science & Technology Translators Journal, 34(1), 31–34, 40. (in Chinese).
- [27] Li, M., & Zhu, X. (2020). Research on machine translation quality assessment: Review and prospects. Chinese Translators Journal, 41(4), 88–95. (in Chinese).
- [28] Hu, K., & Li, Y. (2016). Corpus-based translation studies: Current situation and prospects. Foreign Language Teaching and Research, 48(2), 286–296. (in Chinese).
- [29] U.S. Census Bureau. (2021). American Community Survey: Language use in the United States. U.S. Government Publishing Office.
- [30] Papineni, K., Roukos, S., Ward, T., & Zhu, W. J. (2002). BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (pp. 311–318). Association for Computational Linguistics.
- [31] Rei, R., Stewart, C., Farinha, A. C., & Lavie, A. (2020). COMET: A neural framework for MT evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (pp. 2685–2702). Association for Computational Linguistics.
- [32] Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. OpenAI Technical Report.
- [33] Zhang, Z., & Wang, Y. (2023). Translation education in the age of generative artificial intelligence: Challenges and responses. Foreign Language World, (3), 2–9. (in Chinese).
