Reliability Evaluation of Large Language Models for Social Media Sentiment Annotation: An Empirical Study Based on Model Agreement and Downstream Tasks

Main Article Content

Wenjing Pi
Changxian He

Keywords

large language models (LLMS), sentiment analysis, user-generated content, zero-shot learning

Abstract

Driven by booming social media and user-generated texts, sentiment analysis stands as a core natural language processing task, yet large-scale high-quality data labeling comes with steep costs and practical barriers. This work assesses how dependable large language models are for sentiment tagging, alongside how their labeled outputs shape subsequent classification effects. We build a dataset containing 1,543 Chinese entertainment comment snippets scraped from Bilibili. Under unified prompting, three LLMs—DeepSeek, Qwen and Doubao—produce zero-shot sentiment tags, while 500 sampled entries receive manual annotation to form authoritative benchmark labels. Cohen’s Kappa is adopted to quantify human-model annotation consistency, and TF-IDF features are input to Logistic Regression, Linear SVM and Random Forest for downstream classification evaluation. Among the three models, Doubao achieves the highest human-label consistency with a κ value of 0.800, exceeding Qwen (κ=0.746) and DeepSeek (κ=0.636). Classifiers trained on Doubao’s labels obtain optimal Macro-F1 values of 0.8125 (SVM) and 0.8174 (Random Forest). Obvious performance discrepancies exist across LLMs; high-quality annotations significantly boost downstream classification accuracy, which highlights the importance of selecting competent LLMs for sentiment labeling tasks.

Abstract 33 | PDF Downloads 11

References

  • [1] Li X, Li J, Wu Y. A global optimization approach to multi-polarity sentiment analysis. PLoS One. 2015 Apr 24;10(4):e0124672.
  • [2] Zhu L, Xu M, Bao Y, Xu Y, Kong X. Deep learning for aspect-based sentiment analysis: a review. PeerJ Comput Sci. 2022 Jul 19;8:e1044.
  • [3] Ruan T, Liu Q, Chang Y. Digital media recommendation system design based on user behavior analysis and emotional feature extraction. PLoS One. 2025 May 19;20(5):e0322768.
  • [4] Bilal M, Khan A, Jan S, Musa S, Ali S. Roman Urdu Hate Speech Detection Using Transformer-Based Model for Cyber Security Applications. Sensors (Basel). 2023 Apr 12;23(8):3909.
  • [5] El Koshiry AM, Eliwa EHI, Abd El-Hafeez T, Khairy M. Detecting cyberbullying using deep learning techniques: a pre-trained glove and focal loss technique. PeerJ Comput Sci. 2024 Mar 27;10:e1961.
  • [6] Jin T, Liu J. A text classification method by integrating mobile inverted residual bottleneck convolution networks and capsule networks with adaptive feature channels. Sci Rep. 2025 Jan 5;15(1):855.
  • [7] Luo, Han, Cai, Meng, Cui, Ying, Spread of Misinformation in Social Networks: Analysis Based on Weibo Tweets, Security and Communication Networks, 2021, 7999760, 23 pages, 2021.
  • [8] Satheakeerthy S, Jesudason D, Bahrami B, Bacchi S, Lee YM, Casson R, Sun M, Chan W. Zero-shot LLM-based visual acuity extraction: a pilot study. BMC Ophthalmol. 2025 Jul 1;25(1):359.
  • [9] Chen J, Zhang X, Wang J, Cao T, Gu C, Li Z, Song Y, Yang L, Zhang Z, Zhang Q, Qian D, Li X. Renji endoscopic submucosal dissection video data set for colorectal neoplastic lesions. Sci Data. 2025 Aug 6;12(1):1366.
  • [10] McHugh ML (2012) Interrater reliability: the kappa statistic. Biochem Med (Zagreb) 22(3):276–282
  • [11] Zhou K, Chen W, Sheng W, Hu B, Yang Y, Yang L, Lu H. Explainable AI with fine-tuned large language models for sustainable cultural heritage management. Sci Rep. 2025 Nov 21;15(1):41370.
  • [12] Garcia-Carmona AM, Prieto ML, Puertas E, Beunza JJ. Leveraging Large Language Models for Accurate Retrieval of Patient Information From Medical Reports: Systematic Evaluation Study. JMIR AI. 2025 Jul 3;4:e68776.
  • [13] Huang Y, Lu J, Wu Q, Zhang G. Weakly Supervised Composed Object Re-Identification With Large Models. IEEE Trans Cybern. 2026 Apr 17;PP.
  • [14] Muhammad Umair Ali et al. From vectors to knowledge graphs: A comprehensive analysis of modern retrieval-augmented generation architectures. Computer Science Review. 2026;61:100925.