Clustering Pospay User Reviews Using Rule-Based Hybrid K-Means and TF-IDF

Authors

  • Suci Tamaro Siahaan Universitas Papua
  • Christian Dwi Suhendra University of Papua
  • Lion Ferdinand Marini University of Papua

DOI:

10.33395/sinkron.v10i4.16556

Keywords:

Text mining, Rule-Based Hybrid K-Means, TF-IDF Bigram, User Reviews, Pospay

Abstract

Pospay is a digital financial services application created by PT Pos Indonesia, which offers a variety of financial transaction services. The huge volume of unstructured data and the exponential growth of the number of users’ assessments in Google Play Store makes it impossible to be analyzed manually. This research aims to classify Pospay customer evaluation to identify the main problems experienced by the user and provide recommendation in upgrading application services. This work adopts the Knowledge Discovery in Databases (KDD) technique. The inputs include 11000 user evaluations, of which 10943 are maintained after pre-processing. The input text has been vectorized using TF-IDF Bigram and the ideal number of clusters has been calculated using the Elbow Method and Silhouette Score. Then, a Rule-Based Hybrid K-Means technique was used, which integrated K-Means++ clustering with rule-based refinement to enhance the interpretability of the clusters. The findings produced five primary clusters that are related to verification and identity, system and error, login and account, transaction and balance, and positive reviews. Authentication and identification and transaction and balance were the most talked-about issues among users from these countries, accounting for 31.3% and 26.8% of discussions respectively. PCA visualization and word cloud analysis further supported the interpretation of each cluster. Overall, the proposed approach effectively grouped user reviews into meaningful topics and can assist developers in identifying service priorities to improve the quality and reliability of the Pospay application.

GS Cited Analysis

Downloads

Download data is not yet available.

References

Annas, M., & Wahab, S. N. (2023). Data Mining Methods: K-Means Clustering Algorithms. International Journal of Cyber and IT Service Management (IJCITSM), 3(1), 40–47. https://iiast.iaic-publisher.org/ijcitsm/index.php/IJCITSM/article/view/122

Dąbrowski, J., Letier, E., Perini, A., & Susi, A. (2022). Analysing app reviews for software engineering: a systematic literature review. Empirical Software Engineering, 27(2). https://doi.org/10.1007/s10664-021-10065-7

Handayani, F. D., & Rosyida, I. (2023). Clustering Review Pengguna Aplikasi Zenius pada Layanan Google Play Store Menggunakan Metode DBSCAN dan HDBSCAN. Emerging Statistics and Data Science Journal, 1(2), 178–191.

June, V. N., Widyawati, F., Dawod, A. Y., & Santoso, H. A. (2025). K-Means Clustering Optimization of Toddler Malnutrition Status Using Elbow Method. Journal of Informatics and Web Engineering EISSN:, 4(2).

Khairuna, R., Nurdin, & Ar Razi. (2025). Analisis Perbandingan Algoritma K-Means Dan K-Medoids Untuk Klusterisasi Teks Ulasan Pada Aplikasi Cookpad. Rabit : Jurnal Teknologi Dan Sistem Informasi Univrab, 10(2), 445–458. https://doi.org/10.36341/rabit.v10i2.6131

Kim, G.-Y., & Han, S. (2022). User Review Analysis of English Learning Applications on Google Play Store Using Text-Mining. Journal of Digital Contents Society, 23(10), 1901–1908. https://doi.org/10.9728/dcs.2022.23.10.1901

Mahendra, K., Rahman, Z. S., Supendar, H., & Fahlapi, R. (2026). Pemetaan Tema Keluhan Pengguna Aplikasi Cek Bansos Menggunakan K- Means dan TF-IDF Berbasis Ulasan Google Play Store. JOURNAL KOMPUTER TERKNOLOGI INFORMASI SISTEM KOMPUTER(JUKTISI9), 5(1), 550–559.

Memon, Z. A., Munawar, N., & Kamal, M. (2023). App store mining for feature extraction: analyzing user reviews. Acta Scientiarum - Technology, 46, 1–16. https://doi.org/10.4025/actascitechnol.v46i1.62867

Mohammed, A. A., Sumari, P., & Attabi, K. (2024). Hybrid K-means and Principal Component Analysis (PCA) for Diabetes Prediction: International Journal of Computing and Digital Systems, 15(1), 1719–1728. https://doi.org/10.12785/ijcds/1501121

Mukti, B. P., Hariguna, T., & Tahyudin, I. (2025). Model Klastering Hybrid Menggunakan Inisialisasi K-means++ dan Algoritma Optimasi Grey Wolf. Jurnal Sistem Dan Teknologi Informasi, 13(2), 286–298. https://doi.org/10.26418/justin.v13i2.88211

Nasser, F. K., & Behadili, S. F. (2022). A Review of Data Mining and Knowledge Discovery Approaches for Bioinformatics ةيجهلهيابلا تامهمعملا لاجم يف قئاقحلا فاشتكاو تانايبلا نيدعت ةقيرطل ضا رعتسا. Iraqi Journal of Science, 63(7), 3169–3188. https://doi.org/10.24996/ijs.2022.63.7.37

Nunkaew, W., Ahmed, M., & Nakfon, P. (2026). A Machine Learning-Based Hybrid K-means and Fuzzy Inference System with Rule-Based Fine-Tuning for Sustainable Supplier Segmentation. In AICCC 2025 - 2025 8th Artificial Intelligence and Cloud Computing Conference (Vol. 1, Number 1). Association for Computing Machinery. https://doi.org/10.1145/3789982.3789986

Nurul, K., Djati, I., & Faiza, N. (2023). Identifying Improvement Strategic from User Application Reviews Group Using K-Means Clustering and TF-IDF Weighting. International Journal of Artificial Intelegence Research, 7(2), 152–159.

Onumanyi, A. J., Molokomme, D. N., Isaac, S. J., & Abu-mahfouz, A. M. (2022). AutoElbow : An Automatic Elbow Detection Method for Estimating the Number of Clusters in a Dataset. Applied Sciences, 12(15), 7515. https://doi.org/https://doi.org/10.3390/app12157515

Pamput, J. P., Muthmainnah, A. R., Risal, A. A. N., & Surianto, D. F. (2025). K-Means++ and TF-IDF for Grouping Library Books by Topic. Paradigma - Jurnal Komputer Dan Informatika, 27(2), 74–82. https://doi.org/10.31294/p.v27i2.8272

Pamungkas, M. D., & Februariyanti, H. (2022). Penerapan Algoritma K-Means Clustering Untuk Mengelompokan Data Review Barang Pada E-Commerce Lazada. SemanTIK, 8(2), 99. https://doi.org/10.55679/semantik.v8i2.29058

Prasetyadi, A., Nugroho, B., & Tohari, A. (2022). A Hybrid K-Means Hierarchical Algorithm for Natural Disaster Mitigation Clustering. Journal of Information and Communication Technology, 2(2), 175–200. https://doi.org/https://doi.org/10.32890/jict2022.21.2.2

Ramadhini, N., Annisaa, S. A., Putri, A. S., Tania, K. D., & Rifai, A. (2026). KLASTERISASI PORTAL BERITA ONLINE MENGGUNAKAN ALGORITMA K-MEANS DENGAN TEKNIK KNOWLEDGE DISCOVERY IN DATABASE. JATI (Jurnal Mahasiswa Teknik Informatika), 10(3), 4079–4087. https://doi.org/https://doi.org/10.36040/jati.v10i3.18100

Sudrajat, R., Hadiana, A. I., & Melina. (2025). Evaluasi Kualitas Klaster Wilayah Rawan Bencana Menggunakan K- Means dengan Silhouette dan Elbow Method. Jurnal Algoritma, 22(2), 127–139. https://doi.org/10.33364/algoritma/v.22-2.2379

Ulummuddin, I., Sari, A. P., & Swari, M. H. P. (2024). PEMANFAATAN DATA ULASAN PENGGUNA UNTUK MEMBANGUN SISTEM KLASTERISASI BERDASARKAN PAIN POINTS MENGGUNAKAN ALGORITMA K-MEANS. Jurnal Teknologi Terpadu, 10(1), 70–76. https://doi.org/https://doi.org/10.54914/jtt.v10i1.1252

Vania, P., & Sari, B. N. (2023). Perbandingan Metode Elbow dan Silhouette untuk Penentuan Jumlah Klaster yang Optimal pada Clustering Produksi Padi menggunakan Algoritma K-Means. Jurnal Ilmiah Wahana Pendidikan, 9(21), 547–558. https://doi.org/https://doi.org/10.5281/zenodo.10081332

Zhu, J., Huang, S., Shi, Y., Wu, K., & Wang, Y. (2022). A Method of K-Means Clustering Based on TF-IDF for Software Requirements Documents Written in Chinese Language. IEICE Transactions on Information and Systems, 105(4), 736–754. https://doi.org/10.1587/transinf.2021EDP7144

Downloads


Crossmark Updates

How to Cite

Siahaan, S. T., Suhendra, C. D. ., & Marini, L. F. . (2026). Clustering Pospay User Reviews Using Rule-Based Hybrid K-Means and TF-IDF. Sinkron : Jurnal Dan Penelitian Teknik Informatika, 10(4), 2424-2435. https://doi.org/10.33395/sinkron.v10i4.16556