Depression Detection on Indonesian Social Media Using Fine-Tuned IndoBERT and SVM
DOI:
10.33395/sinkron.v10i3.16294Keywords:
Depression detection, IndoBERT, Indonesian language, Social media, Support vector machineAbstract
Depression has become a major mental health issue in Indonesia, where approximately 167 million of the country’s 273 million citizens actively use social media platforms such as X (Twitter). The informal writing style, code-mixing, and linguistic variability in Indonesian tweets create significant challenges for automated depression detection systems. This study evaluates a fine-tuned IndoBERT model combined with a Support Vector Machine (SVM) classifier for detecting depression-related indications from Indonesian-language tweets. A total of 10,082 Indonesian tweets were collected and labeled into two categories: Terindikasi Depresi and Tidak Terindikasi; after deduplication, 3,874 unique tweets were used for modeling. Two scenarios were compared: (1) a fine-tuned IndoBERT model, and (2) fine-tuned IndoBERT CLS embeddings with a linear SVM classifier. The fine-tuned IndoBERT model achieved 73.20% accuracy (AUC-ROC = 0.8212), while the hybrid approach achieved a marginally higher 73.71% accuracy (AUC-ROC = 0.8088); a McNemar’s test found this difference not statistically significant (p = 0.86). Both models outperformed five traditional TF-IDF-based baselines (best: 70.36%) on the same held-out test set. The hybrid model required only 0.01 MB of storage versus 475.24 MB for the full fine-tuned model. Given statistically equivalent accuracy, combining fine-tuned IndoBERT embeddings with SVM offers substantially lower storage requirements at no measurable cost in classification performance, making it a promising, resource-efficient approach for depression detection on Indonesian social media.
Downloads
References
Amanat, A., Rizwan, M., Javed, A. R., Abdelhaq, M., Alsaqour, R., Pandya, S., & Uddin, M. (2022). Deep Learning for Depression Detection from Textual Data. Electronics, 11(5), 676. https://doi.org/10.3390/electronics11050676
Balasaranya, K., & Ezhumalai, P. (2026). Text Optimized XLNet Based Sentimental Analysis for Opinion Mining Towards Products From Twitter Data. International Journal of Computational Intelligence Systems, 19(1), 165. https://doi.org/10.1007/s44196-026-01265-4
Baydili, İ., Tasci, B., & Tasci, G. (2025). Deep Learning-Based Detection of Depression and Suicidal Tendencies in Social Media Data with Feature Selection. Behavioral Sciences, 15(3), 352. https://doi.org/10.3390/bs15030352
Bendebane, L., Laboudi, Z., Saighi, A., Al-Tarawneh, H., Ouannas, A., & Grassi, G. (2023). A Multi-Class Deep Learning Approach for Early Detection of Depressive and Anxiety Disorders Using Twitter Data. Algorithms, 16(12), 543. https://doi.org/10.3390/a16120543
Bokolo, B. G., & Liu, Q. (2023). Deep Learning-Based Depression Detection from Social Media: Comparative Evaluation of ML and Transformer Techniques. Electronics, 12(21), 4396. https://doi.org/10.3390/electronics12214396
Budiman, A. (2021). DatasetIndikasiDepresi. GitHub Repository. https://github.com/andrebudiman/DatasetIndikasiDepresi.
Cahyawijaya, S., Winata, G. I., Wilie, B., Vincentio, K., Li, X., Kuncoro, A., … Fung, P. (2021). IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Language Generation. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 8875–8898. Stroudsburg, PA, USA: Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.emnlp-main.699
Darmawan, F., Joe, M., Kurniawan, Y. I., & Afuan, L. (2023). Analisis Sentimen Kemungkinan Depresi dan Kecemasan pada Twitter Menggunakan Support Vector Machine. Jurnal Eksplora Informatika, 13(1), 24–36. https://doi.org/10.30864/eksplora.v13i1.854
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North, 4171–4186. Stroudsburg, PA, USA: Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423
Guntuku, S. C., Yaden, D. B., Kern, M. L., Ungar, L. H., & Eichstaedt, J. C. (2017). Detecting depression and mental illness on social media: an integrative review. Current Opinion in Behavioral Sciences, 18, 43–49. https://doi.org/10.1016/j.cobeha.2017.07.005
Hidayat, I. R., & Maharani, W. (2022). General Depression Detection Analysis Using IndoBERT Method. International Journal on Information and Communication Technology (IJoICT), 8(1), 41–51. https://doi.org/10.21108/ijoict.v8i1.634
Jude Chukwura Obi. (2023). A comparative study of several classification metrics and their performances on data. World Journal of Advanced Engineering Technology and Sciences, 8(1), 308–314. https://doi.org/10.30574/wjaets.2023.8.1.0054
Koto, F., Lau, J. H., & Baldwin, T. (2021). IndoBERTweet: A Pretrained Language Model for Indonesian Twitter with Effective Domain-Specific Vocabulary Initialization. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 10660–10668. Stroudsburg, PA, USA: Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.emnlp-main.833
Liu, Q., He, H., Yang, J., Feng, X., Zhao, F., & Lyu, J. (2020). Changes in the global burden of depression from 1990 to 2017: Findings from the Global Burden of Disease study. Journal of Psychiatric Research, 126, 134–140. https://doi.org/10.1016/j.jpsychires.2019.08.002
Porwal, A., Saritha, S. K., & Ahirwal, M. K. (2024). A Hybrid Approach for Depression Classification Using BERT and SVM. https://doi.org/10.1007/978-981-97-3180-0_30
Poświata, R., & Perełkiewicz, M. (2022). Detecting Signs of Depression from Social Media Text using RoBERTa Pre-trained Language Models. Proceedings of the Second Workshop on Language Technology for Equality, Diversity and Inclusion, 276–282. Stroudsburg, PA, USA: Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.ltedi-1.40
Qasim, R., Bangyal, W. H., Alqarni, M. A., & Ali Almazroi, A. (2022). A Fine-Tuned BERT-Based Transfer Learning Approach for Text Classification. Journal of Healthcare Engineering, 2022, 1–17. https://doi.org/10.1155/2022/3498123
Rahayu, K., Fitria, V., Septhya, D., Rahmaddeni, R., & Efrizoni, L. (2023). Klasifikasi Teks untuk Mendeteksi Depresi dan Kecemasan pada Pengguna Twitter Berbasis Machine Learning. MALCOM: Indonesian Journal of Machine Learning and Computer Science, 3(2), 108–114. https://doi.org/10.57152/malcom.v3i2.780
Ridha, M., Nurjanah, D., & Rakha, M. (2024). Multilabel Classification Abusive Language and Hate Speech on Indonesian Twitter Using Transformer Model: IndoBERTweet & IndoRoBERTa. 2024 International Conference on Intelligent Cybernetics Technology & Applications (ICICyTA), 48–54. IEEE. https://doi.org/10.1109/ICICYTA64807.2024.10912874
Saadah, S., Kaenova Mahendra Auditama, Ananda Affan Fattahila, Fendi Irfan Amorokhman, Annisa Aditsania, & Aniq Atiqi Rohmawati. (2022). Implementation of BERT, IndoBERT, and CNN-LSTM in Classifying Public Opinion about COVID-19 Vaccine in Indonesia. Jurnal RESTI (Rekayasa Sistem Dan Teknologi Informasi), 6(4), 648–655. https://doi.org/10.29207/resti.v6i4.4215
Saxena, A., & Santhanavijayan, A. (2026). A layer-wise survey on internal modifications in BERT and its variants: techniques, applications, and performance trade-offs. International Journal of Data Science and Analytics, 22(1), 1. https://doi.org/10.1007/s41060-025-00986-7
Segev, E. (2023). Sharing Feelings and User Engagement on Twitter: It’s All About Me and You. Social Media + Society, 9(2). https://doi.org/10.1177/20563051231183430
Sulak, S. A., & Koklu, N. (2024). Analysis of Depression, Anxiety, Stress Scale (DASS‐42) With Methods of Data Mining. European Journal of Education, 59(4). https://doi.org/10.1111/ejed.12778
Tejaswini, V., Sathya Babu, K., & Sahoo, B. (2024). Depression Detection from Social Media Text Analysis using Natural Language Processing Techniques and Hybrid Deep Learning Model. ACM Transactions on Asian and Low-Resource Language Information Processing, 23(1), 1–20. https://doi.org/10.1145/3569580
Tharwat, A. (2021). Classification assessment methods. Applied Computing and Informatics, 17(1), 168–192. https://doi.org/10.1016/j.aci.2018.08.003
Wilie, B., Vincentio, K., Winata, G. I., Cahyawijaya, S., Li, X., Lim, Z. Y., … Purwarianti, A. (2020). IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding. Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, 843–857. Stroudsburg, PA, USA: Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.aacl-main.85
Downloads
How to Cite
Issue
Section
License
Copyright (c) 2026 Donny Amanullah Putra Rahman, Muhamad Akrom, Muhammad Naufal

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.






















Moraref
PKP Index
Indonesia OneSearch
OCLC Worldcat
Index Copernicus
Scilit
