LLM-Assisted Annotation in Comparative Sentiment and Trend Analysis of Indonesian K-12 Education on X
DOI:
10.33395/sinkron.v10i4.16433Keywords:
K-12 Education, Large Language Model, Sentiment Analysis, Social Media, Transformer-Based ModelsAbstract
Indonesia was ranked 69th out of 81 participating nations in the 2022 PISA test results, reflecting persistent challenges in K-12 education quality. Public discourse on these challenges has grown substantially on social media, yet automated sentiment analysis of Indonesian education-related text remains difficult due to informal language and implicit sentiment expression. Prior work on Indonesian education sentiment has focused on short observation windows or a single policy, without comparing transformer-based models directly against traditional classifiers or scaling annotation with large language models. This study aims to compare traditional machine learning classifiers and transformer-based models for sentiment analysis of Indonesian K-12 education discourse on X, as well as the trend of sentiments longitudinally and key terms associated with different sentiments. The dataset comprises 38,957 tweets in the Indonesian language collected between January 2022 and October 2025 that were annotated via a multi-stage, large language model-assisted few-shot classification pipeline with self-consistency sampling tiebreaker, reaching inter-annotator agreement of 82.01% (Cohen’s kappa = 0.672). This study evaluated the accuracy and macro F1 score of four traditional classifiers, including Logistic Regression, Support Vector Machine, Naïve Bayes, and Random Forest, alongside three transformer-based models, namely IndoBERTweet, IndoBERT-base-p1, and IndoBERT-large-p1. Transformer-based models performed better than traditional classifiers based on all metrics. IndoBERTweet reached the best performance with 0.884 accuracy and 0.859 macro F1, compared to Logistic Regression as the best traditional classifier achieved 0.768 and 0.714, respectively. Negative sentiment dominated public discourse, with “kurikulum merdeka” as the most frequent term across all sentiment categories.
Downloads
References
Adhim, N. F., & Cahyono, N. (2025). Optimization of IndoBERT for Sentiment Analysis of FOMO on Social Media Through Fine-Tuning and Hybrid Labeling. Journal of Applied Informatics and Computing (JAIC), 9(6), 3786–3797. https://doi.org/10.30871/jaic.v9i6.11686
Afandi, M. R., Siswanto, & Adelakun, N. O. (2026). Systematic Literature Review: The Role of Education in Reducing Economic Inequality and Poverty. Sinergi International Journal of Education, 4(1), 1–6. https://doi.org/10.61194/education.v4i1.893
Aggarwal, P., Madaan, A., & Yang, Y. (2023). Let’s Sample Step by Step: Adaptive-Consistency for Efficient Reasoning and Coding with LLMs. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 12375–12396. https://doi.org/10.18653/v1/2023.emnlp-main.761
Cahyawijaya, S., Lovenia, H., Singapore, A. I., & Fung, P. (2024). LLMs Are Few-Shot In-Context Low-Resource Language Learners. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 405–433. https://doi.org/10.18653/v1/2024.naacl-long.24
Dou, L., Liu, Q., Zeng, G., Guo, J., Zhou, J., Mao, X., Jin, Z., Lu, W., & Lin, M. (2024). Sailor: Open Language Models for South-East Asia. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 424–435. https://doi.org/10.18653/v1/2024.emnlp-demo.45
Fitriyani, V., Wibowo, F. A., Wulandhari, L. A., & Nabiilah, G. Z. (2025). A Comparison of Computational Approaches for Sentiment Analysis on Public Opinions About Education in Indonesia. Proceedings of the 2025 IEEE International Conference on Industry 4.0, Artificial Intelligence, and Communications Technology, IAICT 2025, 86–93. https://doi.org/10.1109/IAICT65714.2025.11100431
Hidayat, M. T., Suryadi, S., Latifannisa, N., Sari, S. N., & Rino, R. (2025). Evolution of The Education Curriculum in Indonesia. Journal of Innovation in Educational and Cultural Research, 6(2), 381–395. https://doi.org/10.46843/jiecr.v6i2.1312
Hidayaturrahman, & Prawira, I. (2024). Leveraging Zero-Shot Learning in Large Language Models for Sentiment Analysis: A Comparative Study on the Indonesian Language. Proceedings - 6th International Conference on Informatics, Multimedia, Cyber and Information System, ICIMCIS 2024, 614–619. https://doi.org/10.1109/ICIMCIS63449.2024.10956237
Koto, F., Lau, J. H., & Baldwin, T. (2021). INDOBERTWEET: A Pretrained Language Model for Indonesian Twitter with Effective Domain-Specific Vocabulary Initialization. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 10660–10668. https://doi.org/10.18653/v1/2021.emnlp-main.833
Kurnianingrum, D., Wijaya, A., Widjojo, R., Ratnapuri, C. I., Kartawinata, B. R., & Yustian, O. R. (2023). From Tweets to Insights: Analyzing Indonesian Public Sentiments on Education. 2023 International Conference on Informatics, Multimedia, Cyber and Information Systems, ICIMCIS 2023, 644–647. https://doi.org/10.1109/ICIMCIS60089.2023.10349070
Medantoro, G. F. S., & Muljono. (2026). Comparative Analysis of IndoBERT and Classic Machine Learning Models for Sentiment Classification of Education Policy on Social Media X. Journal of Applied Informatics and Computing (JAIC), 10(1), 548. https://doi.org/10.30871/jaic.v10i1.11723
Niimi, J. (2025). A Simple Ensemble Strategy for LLM Inference: Towards More Stable Text Classification. Natural Language Processing and Information Systems, 189–199. https://doi.org/10.1007/978-3-031-97144-0_17
OECD. (2023). PISA 2022 Results (Volume I): The State of Learning and Equity in Education (Vol. 1). OECD Publishing. https://doi.org/https://doi.org/10.1787/53f23881-en
Opitz, J. (2024). A Closer Look at Classification Evaluation Metrics and a Critical Reflection of Common Evaluation Practice. Transactions of the Association for Computational Linguistics, 12, 820. https://doi.org/10.1162/tacl_a_00675
Riyanto, S., Sitanggang, I. S., Djatna, T., & Atikah, T. D. (2023). Comparative Analysis using Various Performance Metrics in Imbalanced Data for Multi-class Text Classification. International Journal of Advanced Computer Science and Applications (IJACSA), 14(6). https://doi.org/10.14569/IJACSA.2023.01406116
Rosani, M., Dian Lestari, N., & Maria Valianti, R. (2025). Transformation of Education to Welcome the Golden Generation of Indonesia 2045. (Jurnal Manajemen, Kepemimpinan, Dan Supervisi Pendidikan), 10(1), 407–427. https://doi.org/10.31851/jmksp.v10i1.18895
Sahabat-AI. (2024). Gemma2 9B CPT Sahabat-AI v1 Instruct. https://huggingface.co/Sahabat-AI/gemma2-9b-cpt-sahabatai-v1-instruct
Sandra, L., Marcel, Gunarso, G., Fredicia, & Riruma, O. W. (2022). Are University Students Independent: Twitter Sentiment Analysis of Independent Learning in Independent Campus Using RoBERTa Base IndoLEM Sentiment Classifier Model. 2021 International Seminar on Machine Learning, Optimization, and Data Science, ISMODE 2021, 249–253. https://doi.org/10.1109/ISMODE53584.2022.9743110
Sidharta, S., Pranoto, H., Gasa, F. M., Kholis, N., & Ong, A. K. S. (2025). Analysis of Public Sentiment on the 17+8 People’s Demands Issue Using IndoBERT and DistilBERT with LLM-Based Data Annotation. 2025 International Conference on Informatics, Multimedia, Cyber and Information System (ICIMCIS), 575–580. https://doi.org/10.1109/icimcis68501.2025.11327025
Smith-Mutegi, D., Mamo, Y., Kim, J., Crompton, H., & McConnell, M. (2025). Perceptions of STEM education and artificial intelligence: a Twitter (X) sentiment analysis. International Journal of STEM Education, 12(1). https://doi.org/10.1186/s40594-025-00527-5
Wafda, A., Fudholi, D. H., & Nugraha, J. (2025). Aspect-Based Sentiment Analysis On Twitter Tweets About The Merdeka Curriculum Using IndoBERT. JITK (Jurnal Ilmu Pengetahuan Dan Teknologi Komputer), 10(3). https://doi.org/10.33480/jitk.v10i3.5692
Wijaya, T. T., Hidayat, W., Hermita, N., Alim, J. A., & Talib, C. A. (2024). Exploring Contributing Factors To PISA 2022 Mathematics Achievement: Insights From Indonesian Teachers. Infinity Journal, 13(1), 139–156. https://doi.org/10.22460/infinity.v13i1.p139-156
Wilie, B., Vincentio, K., Indra Winata, G., Cahyawijaya, S., Li, X., Lim, Z. Y., Soleman, S., Mahendra, R., Fung, P., Bahar, S., Purwarianti, A., & Bandung, I. T. (2020). IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding. Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, 843–857. https://doi.org/10.18653/v1/2020.aacl-main.85
Wu, J., Wang, X., & Jia, W. (2024). Enhancing Text Annotation through Rationale-Driven Collaborative Few-Shot Prompting. http://arxiv.org/abs/2409.09615
Zhang, W., Chan, H. P., Zhao, Y., Aljunied, M., Wang, J., Liu, C., Deng, Y., Hu, Z., Xu, W., Chia, K., Li, X., & Bing, L. (2025). SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations), 96–105. https://doi.org/10.18653/v1/2025.naacl-demo.10
Downloads
How to Cite
Issue
Section
License
Copyright (c) 2026 Abdurrohhim S. Wahyudi, Rachmat Ramadhiansyah, Indra Budi, Prabu Kresna Putra, Aris Budi Santoso

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.






















Moraref
PKP Index
Indonesia OneSearch
OCLC Worldcat
Index Copernicus
Scilit
