Depression Detection on Indonesian Social Media Using Fine-Tuned IndoBERT and SVM

Authors

  • Donny Amanullah Putra Rahman Department of Informatics Engineering, Faculty of Computer Science, Universitas Dian Nuswantoro, Semarang, Indonesia
  • Muhamad Akrom Department of Informatics Engineering, Faculty of Computer Science, Universitas Dian Nuswantoro, Semarang, Indonesia
  • Muhammad Naufal Department of Informatics Engineering, Faculty of Computer Science, Universitas Dian Nuswantoro, Semarang, Indonesia

DOI:

10.33395/sinkron.v10i3.16294

Keywords:

Depression detection, IndoBERT, Indonesian language, Social media, Support vector machine

Abstract

Depression has become a major mental health issue in Indonesia, where approximately 167 million of the country’s 273 million citizens actively use social media platforms such as X (Twitter). The informal writing style, code-mixing, and linguistic variability in Indonesian tweets create significant challenges for automated depression detection systems. This study evaluates a fine-tuned IndoBERT model combined with a Support Vector Machine (SVM) classifier for detecting depression-related indications from Indonesian-language tweets. A total of 10,082 Indonesian tweets were collected and labeled into two categories: Terindikasi Depresi and Tidak Terindikasi; after deduplication, 3,874 unique tweets were used for modeling. Two scenarios were compared: (1) a fine-tuned IndoBERT model, and (2) fine-tuned IndoBERT CLS embeddings with a linear SVM classifier. The fine-tuned IndoBERT model achieved 73.20% accuracy (AUC-ROC = 0.8212), while the hybrid approach achieved a marginally higher 73.71% accuracy (AUC-ROC = 0.8088); a McNemar’s test found this difference not statistically significant (p = 0.86). Both models outperformed five traditional TF-IDF-based baselines (best: 70.36%) on the same held-out test set. The hybrid model required only 0.01 MB of storage versus 475.24 MB for the full fine-tuned model. Given statistically equivalent accuracy, combining fine-tuned IndoBERT embeddings with SVM offers substantially lower storage requirements at no measurable cost in classification performance, making it a promising, resource-efficient approach for depression detection on Indonesian social media.

GS Cited Analysis

Downloads

Download data is not yet available.

References

Amanat, A., Rizwan, M., Javed, A. R., Abdelhaq, M., Alsaqour, R., Pandya, S., & Uddin, M. (2022). Deep Learning for Depression Detection from Textual Data. Electronics, 11(5), 676. https://doi.org/10.3390/electronics11050676

Balasaranya, K., & Ezhumalai, P. (2026). Text Optimized XLNet Based Sentimental Analysis for Opinion Mining Towards Products From Twitter Data. International Journal of Computational Intelligence Systems, 19(1), 165. https://doi.org/10.1007/s44196-026-01265-4

Baydili, İ., Tasci, B., & Tasci, G. (2025). Deep Learning-Based Detection of Depression and Suicidal Tendencies in Social Media Data with Feature Selection. Behavioral Sciences, 15(3), 352. https://doi.org/10.3390/bs15030352

Bendebane, L., Laboudi, Z., Saighi, A., Al-Tarawneh, H., Ouannas, A., & Grassi, G. (2023). A Multi-Class Deep Learning Approach for Early Detection of Depressive and Anxiety Disorders Using Twitter Data. Algorithms, 16(12), 543. https://doi.org/10.3390/a16120543

Bokolo, B. G., & Liu, Q. (2023). Deep Learning-Based Depression Detection from Social Media: Comparative Evaluation of ML and Transformer Techniques. Electronics, 12(21), 4396. https://doi.org/10.3390/electronics12214396

Budiman, A. (2021). DatasetIndikasiDepresi. GitHub Repository. https://github.com/andrebudiman/DatasetIndikasiDepresi.

Cahyawijaya, S., Winata, G. I., Wilie, B., Vincentio, K., Li, X., Kuncoro, A., … Fung, P. (2021). IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Language Generation. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 8875–8898. Stroudsburg, PA, USA: Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.emnlp-main.699

Darmawan, F., Joe, M., Kurniawan, Y. I., & Afuan, L. (2023). Analisis Sentimen Kemungkinan Depresi dan Kecemasan pada Twitter Menggunakan Support Vector Machine. Jurnal Eksplora Informatika, 13(1), 24–36. https://doi.org/10.30864/eksplora.v13i1.854

Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North, 4171–4186. Stroudsburg, PA, USA: Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423

Guntuku, S. C., Yaden, D. B., Kern, M. L., Ungar, L. H., & Eichstaedt, J. C. (2017). Detecting depression and mental illness on social media: an integrative review. Current Opinion in Behavioral Sciences, 18, 43–49. https://doi.org/10.1016/j.cobeha.2017.07.005

Hidayat, I. R., & Maharani, W. (2022). General Depression Detection Analysis Using IndoBERT Method. International Journal on Information and Communication Technology (IJoICT), 8(1), 41–51. https://doi.org/10.21108/ijoict.v8i1.634

Jude Chukwura Obi. (2023). A comparative study of several classification metrics and their performances on data. World Journal of Advanced Engineering Technology and Sciences, 8(1), 308–314. https://doi.org/10.30574/wjaets.2023.8.1.0054

Koto, F., Lau, J. H., & Baldwin, T. (2021). IndoBERTweet: A Pretrained Language Model for Indonesian Twitter with Effective Domain-Specific Vocabulary Initialization. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 10660–10668. Stroudsburg, PA, USA: Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.emnlp-main.833

Liu, Q., He, H., Yang, J., Feng, X., Zhao, F., & Lyu, J. (2020). Changes in the global burden of depression from 1990 to 2017: Findings from the Global Burden of Disease study. Journal of Psychiatric Research, 126, 134–140. https://doi.org/10.1016/j.jpsychires.2019.08.002

Porwal, A., Saritha, S. K., & Ahirwal, M. K. (2024). A Hybrid Approach for Depression Classification Using BERT and SVM. https://doi.org/10.1007/978-981-97-3180-0_30

Poświata, R., & Perełkiewicz, M. (2022). Detecting Signs of Depression from Social Media Text using RoBERTa Pre-trained Language Models. Proceedings of the Second Workshop on Language Technology for Equality, Diversity and Inclusion, 276–282. Stroudsburg, PA, USA: Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.ltedi-1.40

Qasim, R., Bangyal, W. H., Alqarni, M. A., & Ali Almazroi, A. (2022). A Fine-Tuned BERT-Based Transfer Learning Approach for Text Classification. Journal of Healthcare Engineering, 2022, 1–17. https://doi.org/10.1155/2022/3498123

Rahayu, K., Fitria, V., Septhya, D., Rahmaddeni, R., & Efrizoni, L. (2023). Klasifikasi Teks untuk Mendeteksi Depresi dan Kecemasan pada Pengguna Twitter Berbasis Machine Learning. MALCOM: Indonesian Journal of Machine Learning and Computer Science, 3(2), 108–114. https://doi.org/10.57152/malcom.v3i2.780

Ridha, M., Nurjanah, D., & Rakha, M. (2024). Multilabel Classification Abusive Language and Hate Speech on Indonesian Twitter Using Transformer Model: IndoBERTweet & IndoRoBERTa. 2024 International Conference on Intelligent Cybernetics Technology & Applications (ICICyTA), 48–54. IEEE. https://doi.org/10.1109/ICICYTA64807.2024.10912874

Saadah, S., Kaenova Mahendra Auditama, Ananda Affan Fattahila, Fendi Irfan Amorokhman, Annisa Aditsania, & Aniq Atiqi Rohmawati. (2022). Implementation of BERT, IndoBERT, and CNN-LSTM in Classifying Public Opinion about COVID-19 Vaccine in Indonesia. Jurnal RESTI (Rekayasa Sistem Dan Teknologi Informasi), 6(4), 648–655. https://doi.org/10.29207/resti.v6i4.4215

Saxena, A., & Santhanavijayan, A. (2026). A layer-wise survey on internal modifications in BERT and its variants: techniques, applications, and performance trade-offs. International Journal of Data Science and Analytics, 22(1), 1. https://doi.org/10.1007/s41060-025-00986-7

Segev, E. (2023). Sharing Feelings and User Engagement on Twitter: It’s All About Me and You. Social Media + Society, 9(2). https://doi.org/10.1177/20563051231183430

Sulak, S. A., & Koklu, N. (2024). Analysis of Depression, Anxiety, Stress Scale (DASS‐42) With Methods of Data Mining. European Journal of Education, 59(4). https://doi.org/10.1111/ejed.12778

Tejaswini, V., Sathya Babu, K., & Sahoo, B. (2024). Depression Detection from Social Media Text Analysis using Natural Language Processing Techniques and Hybrid Deep Learning Model. ACM Transactions on Asian and Low-Resource Language Information Processing, 23(1), 1–20. https://doi.org/10.1145/3569580

Tharwat, A. (2021). Classification assessment methods. Applied Computing and Informatics, 17(1), 168–192. https://doi.org/10.1016/j.aci.2018.08.003

Wilie, B., Vincentio, K., Winata, G. I., Cahyawijaya, S., Li, X., Lim, Z. Y., … Purwarianti, A. (2020). IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding. Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, 843–857. Stroudsburg, PA, USA: Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.aacl-main.85

Downloads


Crossmark Updates

How to Cite

Rahman, D. A. P., Akrom, M. ., & Naufal, M. . (2026). Depression Detection on Indonesian Social Media Using Fine-Tuned IndoBERT and SVM. Sinkron : Jurnal Dan Penelitian Teknik Informatika, 10(3). https://doi.org/10.33395/sinkron.v10i3.16294

Most read articles by the same author(s)