Implementation of Integrity Zone Document Classification Using IndoBERT Model and Logistic Regression

Authors

  • Adidtya Perdana Computer Science Study Program, Universitas Negeri Medan, Indonesia
  • Nurul Ain Farhana Statistics Study Program, Universitas Negeri Medan, Indonesia
  • Didi Febrian Mathematics Study Program, Universitas Negeri Medan, Indonesia

DOI:

10.33395/sinkron.v10i4.16585

Abstract

The Indonesian government's Integrity Zone program mandates systematic classification of administrative documents into hierarchical compliance codes, yet manual categorization remains labor-intensive, inconsistent, and unscalable for higher education institutions. This study aims to develop and evaluate an automated multi-label classification pipeline that maps Indonesian bureaucratic documents to hierarchical compliance codes while maintaining computational efficiency for institutional deployment. A curated dataset of 330 Integrity Zone documents from the Faculty of Mathematics and Natural Sciences, Universitas Negeri Medan, annotated across 82 hierarchical codes, was processed using a frozen IndoBERT encoder to extract 768-dimensional contextual embeddings. These features were classified using a balanced One-vs-Rest Logistic Regression model, with decision thresholds optimized via grid search on a held-out validation set to balance precision and recall. The complete pipeline was deployed as a REST microservice integrated into an existing PHP-based document management system. On a held-out test set of 66 documents not used for training, validation, or threshold selection, the pipeline achieved a Macro F1-score of 0.872, Micro F1-score of 0.894, Hamming Loss of 0.082, and a samples-averaged accuracy of 0.917 (pooled label-wise accuracy 0.918; subset accuracy 0.412). The optimized decision threshold of 0.38 favored recall over the conventional 0.50 cutoff, consistent with the higher institutional cost of missing a relevant compliance code. End-to-end inference latency averaged 2.36 seconds on the deployment server. The proposed pipeline shows practical viability as a decision-support tool for resource-constrained public institutions operating under mandatory human verification; the present evaluation is nonetheless limited by a small held-out test set, and several methodological details are reported in full in the Method section to support reproducibility and to rule out validation leakage.

GS Cited Analysis

Downloads

Download data is not yet available.

Author Biographies

Adidtya Perdana, Computer Science Study Program, Universitas Negeri Medan, Indonesia

Computer Science

Nurul Ain Farhana, Statistics Study Program, Universitas Negeri Medan, Indonesia

Statistics

Didi Febrian, Mathematics Study Program, Universitas Negeri Medan, Indonesia

Mathematics

References

Alhammad, M., Avdelidis, N. P., Ibarra Castanedo, C., Maldague, X., Zolotas, A., Torbali, E., & Genest, M. (2024). Multi-label classification algorithms for composite materials under infrared thermography testing. Quantitative InfraRed Thermography Journal, 21(1), 3–29. https://doi.org/10.1080/17686733.2022.2126638

Alnuaimi, A. F. A. H., & Albaldawi, T. H. K. (2024). An overview of machine learning classification techniques. BIO Web of Conferences, 97, 00133. https://doi.org/10.1051/BIOCONF/20249700133

Awal Kassim, M., Viktor, H., & Michalowski, W. (2024). Multi-Label Lifelong Machine Learning: A Scoping Review of Algorithms, Techniques, and Applications. IEEE Access, 12, 74539–74557. https://doi.org/10.1109/ACCESS.2024.3403569

Chen, H., Wu, L., Chen, J., Lu, W., & Ding, J. (2022). A comparative study of automated legal text classification using random forests and deep learning. Information Processing & Management, 59(2), 102798. https://doi.org/10.1016/J.IPM.2021.102798

Ding, Y., Han, X., Yang, J., Wang, T., Bi, Z., Song, X., Hao, J., Song, J., Ge, E., Peng, B., Liu, Z., Liang, C. X., Zhang, Y., Liu, M., Xu, J., Huang, B., Mo, Y., Yu, Z., Qiao, J., … Ma, Y. (2026). Cross-Lingual Transfer Learning in Large Language Models: Multilingual Representations and Low-Resource Adaptation. https://doi.org/10.36227/TECHRXIV.176799777.79169627/V1

Endut, N., Amir, W. M., Hamzah, F. W., Ismail, I., Yusof, M. K., Baker, Y. A., & Yusoff, H. (2022). A Systematic Literature Review on Multi-Label Classification based on Machine Learning Algorithms. UIKTEN - Association for Information Communication Technology Education and Science, 11(2), 658–666.

Gardazi, N. M., Daud, A., Malik, M. K., Bukhari, A., Alsahfi, T., & Alshemaimri, B. (2025). BERT applications in natural language processing: a review. Artificial Intelligence Review 2025 58:6, 58(6), 166-. https://doi.org/10.1007/S10462-025-11162-5

Gupta, S., Yadav, A., Yadav, D., & Dixit, U. (2022). Analysis of Automatic Text Classification of Legal Documents. SSRN Electronic Journal. https://doi.org/10.2139/SSRN.4288439

Kirasich, K. ;, Smith, T. ;, & Sadler, B. (2018). Random Forest vs Logistic Regression: Binary Classification for Heterogeneous Datasets. SMU Data Science Review, 1(3), 9. https://scholar.smu.edu/datasciencereviewAvailableat:https://scholar.smu.edu/datasciencereview/vol1/iss3/9http://digitalrepository.smu.edu.

Koto, F., Rahimi, A., Lau, J. H., & Baldwin, T. (2020). IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP. COLING 2020 - 28th International Conference on Computational Linguistics, Proceedings of the Conference, 757–770. https://doi.org/10.18653/V1/2020.COLING-MAIN.66

Kumar, V., Singh, R. S., Rambabu, M., & Dua, Y. (2024). Deep learning for hyperspectral image classification: A survey. Computer Science Review, 53, 100658. https://doi.org/10.1016/J.COSREV.2024.100658

NaulakChingmuankim. (2022). A comparative study of Naive Bayes Classifiers with improved technique on Text Classification. https://doi.org/10.36227/TECHRXIV.19918360.V1

Perdana, A., Dewi, S., Farhana, N. A., & Febrian, D. (2025). Comparative Analysis of SDLC and R&D Methods in System Development: A Case Study of Integrity Zone Management System. Sinkron : Jurnal Dan Penelitian Teknik Informatika, 9(4), 3197–3209. https://doi.org/10.33395/SINKRON.V9I4.15337

Perdana, A., Farhana, N. A., Harliana, P., Muslim, I., & Karo, K. (2024). Web-Based Application Development using PHP-Native Framework on Agent of Change Integrity Zone Information System. Jurnal.Polgan.Ac.IdA Perdana, NA Farhana, P Harliana, IMK KaroSinkron: Jurnal Dan Penelitian Teknik Informatika, 2024•jurnal.Polgan.Ac.Id, 8(4). https://doi.org/10.33395/sinkron.v8i4.14118

Rajamani, S. K., & Iyer, R. S. (2022). Machine Learning-Based Mobile Applications Using Python and Scikit-Learn. Quantitative InfraRed Thermography Journal, 21(1), 282–306. https://doi.org/10.4018/978-1-6684-8582-8.CH016

Shah, K., Patel, H., Sanghvi, D., & Shah, M. (2020). A Comparative Analysis of Logistic Regression, Random Forest and KNN Models for the Text Classification. Augmented Human Research, 5(1). https://doi.org/10.1007/s41133-020-00032-0

Singh, D., Bhatnagar, M., & Yadav, V. (2022). PDF Classification Using Logistic Regression and Latent Dirichlet Allocation. Lecture Notes in Networks and Systems, 237, 399–407. https://doi.org/10.1007/978-981-16-6407-6_36/SAVE-RESEARCH

Tarekegn, A. N., Ullah, M., & Cheikh, F. A. (2024). Deep Learning for Multi-Label Learning: A Comprehensive Survey. https://arxiv.org/pdf/2401.16549

Wang, Y., Sun, Y., Fu, Y., Zhu, D., & Tian, Z. (2024). Spectrum-BERT: Pretraining of Deep Bidirectional Transformers for Spectral Classification of Chinese Liquors. IEEE Transactions on Instrumentation and Measurement, 73, 1–13. https://doi.org/10.1109/TIM.2024.3374300

Wilie, B., Vincentio, K., Winata, G. I., Cahyawijaya, S., Li, X., Lim, Z. Y., Soleman, S., Mahendra, R., Fung, P., Bahar, S., & Purwarianti, A. (2020). IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding. Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, AACL-IJCNLP 2020, 843–857. https://doi.org/10.18653/v1/2020.aacl-main.85

Downloads


Crossmark Updates

How to Cite

Perdana, A. ., Farhana, N. A., & Febrian, D. (2026). Implementation of Integrity Zone Document Classification Using IndoBERT Model and Logistic Regression. Sinkron : Jurnal Dan Penelitian Teknik Informatika, 10(4). https://doi.org/10.33395/sinkron.v10i4.16585

Most read articles by the same author(s)