Comparison of Naive Bayes and Support Vector Machine for Sentiment Analysis of BPJS Health Service Deactivation

Authors

  • Vitriayanti Payung Allo University of Papua
  • Marlinda Sanglise University of Papua, Indonesia
  • Julius Panda Putra Naibaho University of Papua, Indonesia

DOI:

10.33395/sinkron.v10i3.16261

Keywords:

5-Fold Cross Validation, BPJS Health, Naive Bayes, Sentiment Analysis, Social Media X, Support Vector Machine, TF-IDF

Abstract

The deactivation of BPJS Health services has become a topic of public discussion on social media, particularly on platform X, where users frequently express opinions regarding healthcare access and membership status. This study aims to analyze public sentiment toward the deactivation of BPJS Health services and compare the performance of Naive Bayes and Support Vector Machine (SVM) for sentiment classification. The dataset consisted of 8,357 tweets collected from social media X, of which 7,546 tweets were retained after preprocessing, including data cleaning, case folding, tokenizing, stopword removal, and stemming. TF-IDF and FastText were employed as text representation techniques, while model evaluation was conducted using 5-Fold Cross Validation, Grid Search Cross Validation for hyperparameter optimization, and a paired t-test for statistical significance analysis. Classification performance was measured using accuracy, precision, recall, and F1-score metrics. The results showed that negative sentiment dominated public opinion, accounting for 70.63% of the dataset, followed by neutral sentiment (26.48%) and positive sentiment (2.89%). The SVM model with TF-IDF achieved the highest performance, with an accuracy of 80.97%, precision of 80.16%, recall of 80.97%, and F1-score of 79.70%, outperforming Naive Bayes with TF-IDF (79.01%), Naive Bayes with FastText (64.26%), and SVM with FastText (80.31%). Furthermore, a paired t-test confirmed that the performance difference between Naive Bayes and SVM was statistically significant (p = 0.011). These findings indicate that SVM combined with TF-IDF is more effective for sentiment classification of high-dimensional social media text data and provide empirical evidence regarding the effectiveness of different text representation approaches for healthcare policy-related sentiment analysis.

 

GS Cited Analysis

Downloads

Download data is not yet available.

Author Biographies

Marlinda Sanglise, University of Papua, Indonesia

Lecturer at the University of Papua.

Julius Panda Putra Naibaho, University of Papua, Indonesia

Lecturer at the University of Papua.

References

A. Fauzi, & D. Nugroho. (2021). Analisis Sentimen Media Sosial Menggunakan Machine Learning. Jurnal Sistem Informasi, 11(2), 98–107.

A. Wibowo, & R. Pratama. (2023). Implementasi TF-IDF dan Support Vector Machine pada Analisis Sentimen Media Sosial Twitter. Jurnal RESTI (Rekayasa Sistem Dan Teknologi Informasi), 7(1), 120–128.

Bojanowski, P., Grave, E., Joulin, A., & Mikolov, T. (2017). Enriching Word Vectors with Subword Information. Transactions of the Association for Computational Linguistics, 5, 135–146.

https://doi.org/10.1162/tacl_a_00051

C. A. Nurhaliza Agustina, R. Novita, Mustakim, & N. E. Rozanda. (2024). The Implementation of TF-IDF and Word2Vec on Booster Vaccine Sentiment Analysis Using Support Vector Machine Algorithm. Procedia Computer Science, 234, 156–163.

D. Iskandar, & Y. Nataliani. (2021). Perbandingan Naive Bayes, SVM, dan K-NN untuk Analisis Sentimen Berbasis Aspek. Jurnal RESTI (Rekayasa Sistem Dan Teknologi Informasi), 5(6), 1120–1126.

E. Elgeldawi, A. Sayed, A. R. Galal, & A. M. Zaki. (2021). Hyperparameter Tuning for Machine Learning Algorithms Used for Arabic Sentiment Analysis. Informatics, 8(4).

F. Z. Tala. (2003). A Study of Stemming Effects on Information Retrieval in Bahasa Indonesia. University of Amsterdam.

Salton, G., & Buckley, C. (1988). Term-Weighting Approaches in Automatic Text Retrieval. Information Processing & Management, 24(5), 513–523.

https://doi.org/10.1016/0306-4573(88)90021-0

Indra, K., Apri Lia, H., Shofa Shofia, H., Agustia, H., Bayu, P., & Aviv Yuniar, R. (2023). Perbandingan Algoritma Naive Bayes Dan SVM Dalam Sentimen Analisis Marketplace Pada Twitter. Jurnal Teknik Informatika Dan Sistem Informasi, 10(1). Retrieved from http://jurnal.mdp.ac.id

Jiawei Han, Micheline Kamber, & Jian Pei. (2012). Data Mining: Concepts and Techniques (3rd Edition). USA: Morgan Kaufmann.

Joachims, T. (1998). Text Categorization with Support Vector Machines: Learning with Many Relevant Features. Proceedings of ECML-98, 137–142.

Kohavi, R. (1995). A Study of Cross-Validation and Bootstrap for Accuracy Estimation and Model Selection. Proceedings of the 14th International Joint Conference on Artificial Intelligence (IJCAI), 1137–1143.

M. Ridwan, & F. Rizki. (2022). Klasifikasi Sentimen Twitter Bahasa Indonesia Menggunakan TF-IDF dan Support Vector Machine. Jurnal Informatika, 8(3), 210–219.

Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient Estimation of Word Representations in Vector Space. arXiv preprint arXiv:1301.3781.P. S. Ghatora, S. E. Hosseini, S. Pervez, M

J. Iqbal, & N. Shaukat. (2024). Sentiment Analysis of Product Reviews Using Machine Learning and Pre-Trained LLM. Big Data and Cognitive Computing, 8(12), 199.

Q. Li, & et al. (2022). A Survey on Text Classification: From Traditional to Deep Learning. ACM Transactions on Intelligent Systems and Technology, 13(2), 31.

R. Amandasari, & N. Damayanti. (2022). Analisis Sentimen Pelayanan BPJS Kesehatan Menggunakan Metode Support Vector Machine dan Naive Bayes. Jurnal Media Informatika Budidarma, 6(3), 1455–1463.

S. Akuma, T. Lubem, & I. T. Adom. (2022). Comparing Bag of Words and TF-IDF with Different Models for Hate Speech Detection from Live Tweets. International Journal of Information Technology, 14(7), 3629–3635.

saddam, Mohd. A., D, E. K., & Indra. (2023). Analisis Sentimen Fenomena PHK Massal Menggunakan Naive Bayes dan Support Vector Machine, 8(3).

Sebastiani, F. (2002). Machine Learning in Automated Text Categorization. Retrieved from www.ira.uka.de/bibliography/Ai/automated.text.

Sholekhah, A., & Muntahanah, M. (2025). Perbandingan Naïve Bayes dan Support Vector Machine Dalam Analisa Sentimen Tentang Penyitaan Aset Koruptor di Twitter. MALCOM: Indonesian Journal of Machine Learning and Computer Science, 5(3), 981–989. doi:10.57152/malcom.v5i3.2068

Trevor Hastie, Robert Tibshirani, & Jerome Friedman. (2009). The Elements of Statistical Learning (2nd Edition). New York: Springer.

Y. Mao, Q. Liu, & Y. Zhang. (2024). Sentiment analysis methods, applications, and challenges: A systematic literature review. Journal of King Saud University - Computer and Information Sciences, 36(4), 102048.

Downloads


Crossmark Updates

How to Cite

Allo, V. P. ., Sanglise, M., & Naibaho, J. P. P. (2026). Comparison of Naive Bayes and Support Vector Machine for Sentiment Analysis of BPJS Health Service Deactivation. Sinkron : Jurnal Dan Penelitian Teknik Informatika, 10(3), 1649-1660. https://doi.org/10.33395/sinkron.v10i3.16261