CHCF-Bench: Safety Alignment Evaluation of Large Language Models for Chinese High-Context Cultural Friction
Main Article Content
Keywords
large language models, cultural alignment, high-context culture, Chinese cultural friction, jailbreak attack, safety evaluation
Abstract
Large language models (LLMs) are increasingly used in cross-cultural communication, content generation, intelligent question answering, and social collaboration. However, existing safety evaluations mainly focus on explicit risks such as illegality, violence, privacy leakage, hate speech, and malicious code, while paying less attention to implicit social harm in high-context cultures. This paper proposes CHCF-Bench, a Chinese High- context Cultural Friction Benchmark containing 100 samples across four scenarios: workplace and power, clan and intergenerational relations, favor and reciprocity, and traditional rituals. CHCF -Bench evaluates whether LLMs can identify potential harm from relational identity, implicit norms, face concerns, reciprocity pressure, and ritual constraints, while producing low-harm responses that balance legal rights, individual dignity, and relational sensitivity. Experiments are conducted with three attack-prompt methods, ICA, DeepInception, and CAPAIR, on six representative LLMs: DeepSeek -V3, Gemini -2.0-Flash, GPT -4o, Meta -Llama-3.1-8B- Instruct, ERNIE -4.0-8K, and GLM -4. Qwen2.5 -72B-Instruct is used as a unified judge. Results show that GLM-4 and Gemini -2.0-Flash reach average a ttack success rates of 79.00% and 78.00%, respectively. DeepInception is the strongest attack method, with an average ASR of 70.00%. Traditional ritual and clan/intergenerational scenarios are most likely to induce high-harm outputs. These findings reveal remaining weaknesses of LLM safety alignment under implicit cultural and relational risk.
References
- [1] Hall, E. T. (1976). Beyond culture. Anchor Press.
- [2] Hofstede, G. (2001). Culture's consequences: Comparing values, behaviors, institutions and organizations across nations (2nd ed.). Sage.
- [3] Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., et al. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35. https://arxiv.org/abs/2203.02155
- [4] Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., & Finn, C. (2023). Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36. https://arxiv.org/abs/2305.18290
- [5] Gehman, S., Gururangan, S., Sap, M., Choi, Y., & Smith, N. A. (2020). RealToxicityPrompts: Evaluating neural toxic degeneration in language models. Findings of the Association for Computational Linguistics: EMNLP 2020, 3356-3369. https://doi.org/10.18653/v1/2020.findings-emnlp.301
- [6] Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P.-S., et al. (2021). Ethical and social risks of harm from language models. arXiv. https://arxiv.org/abs/2112.04359
- [7] Gabriel, I. (2020). Artificial intelligence, values, and alignment. Minds and Machines, 30, 411 -437. https://doi.org/10.1007/s11023-020-09539-2
- [8] Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., et al. (2021). On the opportunities and risks of foundation models. arXiv. https://arxiv.org/abs/2108.07258
- [9] Ji, J., Qiu, T., Chen, B., Zhang, B., Lou, H., Wang, K., et al. (2023). AI alignment: A comprehensive survey. arXiv. https://arxiv.org/abs/2310.19852
- [10] Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., et al. (2022). Constitutional AI: Harmlessness from AI feedback. arXiv. https://arxiv.org/abs/2212.08073
- [11] Hendrycks, D., Burns, C., Basart, S., Critch, A., Li, J., Song, D., & Steinhardt, J. (2021). Aligning AI with shared human values. International Conference on Learning Representations. https://arxiv.org/abs/2008.02275
- [12] Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., & Fredrikson, M. (2023). Universal and transferable adversarial attacks on aligned language models. arXiv. https://arxiv.org/abs/2307.15043
- [13] Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., & Wong, E. (2023). Jailbreaking black box large language models in twenty queries. arXiv. https://arxiv.org/abs/2310.08419
- [14] Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., et al. (2024). HarmBench: A standardized evaluation framework for automated red teaming and robust refusal. arXiv. https://arxiv.org/abs/2402.04249
- [15] Zheng, L., Chiang, W. L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., et al. (2023). Judging LLM-as-a- judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems, 36. https://arxiv.org/abs/2306.05685
- [16] Russinovich, M., Salem, A., & Eldan, R. (2024). Great, now write an article about that: The Crescendo multi-turn LLM jailbreak attack. arXiv. https://arxiv.org/abs/2404.01833
- [17] Li, H., Ye, J., Wu, J., Yan, T., Wang, C., & Li, Z. (2024). JailPO: A novel black-box jailbreak framework via preference optimization against aligned LLMs. arXiv. https://arxiv.org/abs/2412.15623
- [18] Li, W., Zhu, L., Song, Y., Lin, R., Mao, R., & You, Y. (2024). Can a large language model be a gaslighter? arXiv. https://arxiv.org/abs/2410.09181
- [19] Yang, Y., Hui, B., Yuan, H., Gong, N., & Cao, Y. (2024). SneakyPrompt: Jailbreaking text-to-image generative models. IEEE Symposium on Security and Privacy. https://arxiv.org/abs/2305.12082
- [20] Huang, Y., Liang, L., Li, T., Jia, X., Wang, R., Miao, W., Pu, G., & Liu, Y. (2024). Perception-guided jailbreak against text-to-image models. arXiv. https://arxiv.org/abs/2408.10848
- [21] Google AI for Developers. (2026). Models: Gemini API. Google. Retrieved July 9, 2026, from https://ai.google.dev/gemini-api/docs/models
- [22] Qwen Team. (2024). Qwen2.5 technical report. arXiv. https://arxiv.org/abs/2412.15115
