Comparative Evaluation of IndoBERT-Based Architectures for Imbalanced Indonesian News Title Classification

Authors

  • Erwin Sirait Department of Computerized Accounting, Politeknik Bisnis Indonesia, Simalungun, Indonesia
  • Juni Ismail Department of Computer Engineering, Politeknik Bisnis Indonesia, Simalungun, Indonesia

DOI:

10.33395/sinkron.v10i3.16209

Keywords:

Attention Pooling, Back-Translation, Focal Loss, Imbalanced Learning, IndoBERT, Indonesian News Classification

Abstract

Stacking multiple imbalance-mitigation techniques on top of a pretrained transformer is widely assumed to compound their individual benefits, yet rigorous component-wise evidence for this assumption remains scarce in the Indonesian text classification literature. Four classification architectures are compared in this work on a publicly available Indonesian news title corpus. The working set contains 27,266 short headlines, drawn as a 30% stratified subsample from a cleaned corpus of 90,891 headlines, spread over nine target categories with a class ratio of 13.29. Three reference architectures are constructed: an LSTM trained from scratch with Random Oversampling, a bidirectional LSTM augmented with additive attention, and a fine-tuned IndoBERT on the oversampled training partition. A fourth architecture extends IndoBERT through three additions, namely learned attention pooling over contextual token embeddings, focal modulation applied on top of the cross-entropy term, and minority-class paraphrasing via Indonesian–English–Indonesian back-translation. Every configuration is evaluated through stratified 5-fold cross-validation, paired t-tests with Bonferroni correction across three comparisons, and McNemar tests on the held-out partition. The fine-tuned IndoBERT with Random Oversampling alone reaches the highest macro F1 of 0.837. By contrast, the combined configuration drops to 0.799, and statistical verification confirms that the gap is systematic rather than attributable to fold-level variation. A component-wise ablation isolates focal modulation as the principal driver of the decline, because it disturbs an already-balanced training distribution. The principal outcome of this study is empirical evidence indicating that composing several imbalance-oriented techniques on a pretrained transformer can yield adverse interactions rather than cumulative gains.

GS Cited Analysis

Downloads

Download data is not yet available.

References

Asrawi, H., Utami, E., & Yaqin, A. (2023). LSTM and bidirectional GRU comparison for text classification. Sinkron: Jurnal dan Penelitian Teknik Informatika, 8(4), 2264–2274. https://doi.org/10.33395/sinkron.v8i4.12899

Fanani, A. M., & Wahyuddin, M. I. (2026). Sarcasm detection in Indonesian YouTube comments using fine-tuned IndoBERT with class imbalance handling. Sinkron: Jurnal dan Penelitian Teknik Informatika, 10(1). https://doi.org/10.33395/sinkron.v10i1.15607

Fathin, M. A., Sibaroni, Y., & Prasetyowati, S. S. (2024). Handling imbalance dataset on hoax Indonesian political news classification using IndoBERT and random sampling. Jurnal Media Informatika Budidarma, 8(1), 352–360. https://doi.org/10.30865/mib.v8i1.7099

Idris, M., Rifai, A., & Tania, K. D. (2025). Sentiment analysis of Tokopedia app reviews using machine learning and word embeddings. Sinkron: Jurnal dan Penelitian Teknik Informatika, 9(1), 210–219. https://doi.org/10.33395/sinkron.v9i1.14278

Khairani, U., Mutiawani, V., & Ahmadian, H. (2024). Pengaruh tahapan preprocessing terhadap model IndoBERT dan IndoBERTweet untuk mendeteksi emosi pada komentar akun berita Instagram. Jurnal Teknologi Informasi dan Ilmu Komputer, 11(4), 887–894. https://doi.org/10.25126/jtiik.1148315

Komisarenko, V., & Kull, M. (2024). Improving calibration by relating focal loss, temperature scaling, and properness. In Proceedings of the 27th European Conference on Artificial Intelligence (ECAI 2024). https://doi.org/10.3233/FAIA240658

Koto, F., Rahimi, A., Lau, J. H., & Baldwin, T. (2020). IndoLEM and IndoBERT: A benchmark dataset and pre-trained language model for Indonesian NLP. In Proceedings of the 28th International Conference on Computational Linguistics (pp. 757–770). International Committee on Computational Linguistics. https://doi.org/10.18653/v1/2020.coling-main.66

Kunaefi, A., Abidin, Z., & Kusumawati, R. (2025). Klasifikasi berita hoaks bahasa Indonesia menggunakan IndoBERT fine-tuning dengan pendekatan focal loss pada data tidak seimbang. JIPI (Jurnal Ilmiah Penelitian dan Pembelajaran Informatika), 10(2), 1706–1714. https://doi.org/10.29100/jipi.v10i2.7811

Lin, T.-Y., Goyal, P., Girshick, R., He, K., & Dollár, P. (2017). Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision (ICCV 2017) (pp. 2980–2988). https://doi.org/10.1109/ICCV.2017.324

Mandhasiya, F. R., Murfi, H., & Bustamam, A. (2024). A hybrid BERT and deep learning model for Indonesian sentiment classification. Information, 15(11), 720. https://doi.org/10.3390/info15110720

Mukhoti, J., Kulharia, V., Sanyal, A., Golodetz, S., Torr, P., & Dokania, P. (2020). Calibrating deep neural networks using focal loss. In Advances in Neural Information Processing Systems (Vol. 33, pp. 15288–15299). Curran Associates, Inc.

Nimasari, A., Saraswati, G. W., & Lutfina, E. (2026). Improving multi-class public complaint classification with LSTM, Word2Vec, and random oversampling. Sinkron: Jurnal dan Penelitian Teknik Informatika, 10(2), 1094–1103. https://doi.org/10.33395/sinkron.v10i2.15975

Rahma, I. A., & Suadaa, L. H. (2023). Penerapan text augmentation untuk mengatasi data yang tidak seimbang pada klasifikasi teks berbahasa Indonesia. Jurnal Teknologi Informasi dan Ilmu Komputer, 10(6), 1329–1340. https://doi.org/10.25126/jtiik.2023107325

Sennrich, R., Haddow, B., & Birch, A. (2016). Improving neural machine translation models with monolingual data. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Vol. 1, pp. 86–96). Association for Computational Linguistics. https://doi.org/10.18653/v1/P16-1009

Suhartono, D., Majiid, M. R. N., & Fredyan, R. (2024). Towards a sarcasm detection model with enhanced text representation for Bahasa Indonesia. Journal of Big Data, 11(1), 173. https://doi.org/10.1186/s40537-024-01024-2

Taskiran, S. F., Turkoglu, B., Kaya, E., & Asuroglu, T. (2025). A comprehensive evaluation of oversampling techniques for enhancing text classification performance. Scientific Reports, 15(1), 22301. https://doi.org/10.1038/s41598-025-05791-7

Utami, S., Lhaksmana, K. M., & Sibaroni, Y. (2023). Deep learning and imbalance handling on movie review sentiment analysis. Sinkron: Jurnal dan Penelitian Teknik Informatika, 8(3), 1894–1907. https://doi.org/10.33395/sinkron.v8i3.12770

Wongso, W., Setiawan, D. S., Limcorn, S., & Joyoadikusumo, A. (2025). NusaBERT: Teaching IndoBERT to be multilingual and multicultural. In Proceedings of the Second Workshop in South East Asian Language Processing (pp. 10–26). Association for Computational Linguistics.

Wongvorachan, T., He, S., & Bulut, O. (2023). A comparison of undersampling, oversampling, and SMOTE methods for dealing with imbalanced classification in educational data mining. Information, 14(1), 54. https://doi.org/10.3390/info14010054

Zahro, A. C., Alzami, F., Sani, R. R., Fahmi, A., Megantara, R. A., Naufal, M., Azies, H. A., & Iswahyudi, I. (2025). Fairer public complaint classification on LaporGub: Integrating XLM-RoBERTa with focal loss for imbalance data. Sinkron: Jurnal dan Penelitian Teknik Informatika, 9(4), 1850–1862. https://doi.org/10.33395/sinkron.v9i4.15260

Downloads


Crossmark Updates

How to Cite

Sirait, E., & Ismail, J. (2026). Comparative Evaluation of IndoBERT-Based Architectures for Imbalanced Indonesian News Title Classification. Sinkron : Jurnal Dan Penelitian Teknik Informatika, 10(3), 1896-1907. https://doi.org/10.33395/sinkron.v10i3.16209