Evaluating the Sensitivity–Specificity Trade-off in Risk Factor-Based Cervical Cancer Classification Using SMOTE and ACO-Optimized XGBoost

Authors

  • Titin Prihatin Universitas Bina Sarana Informatika
  • Suharjanti Suharjanti Universitas Bina Sarana Informatika
  • Resti Lia Andharsaputri Universitas Bina Sarana Informatika
  • Hafis Nurdin Universitas Bina Sarana Informatika

DOI:

10.33395/sinkron.v10i4.16797

Keywords:

Ant Colony Optimization, Cervical Cancer, Class Imbalance, SMOTE, Nested Cross-Validation, XGBoost

Abstract

Class imbalance is a major challenge in risk factor-based cervical cancer classification because positive cases are substantially less frequent than negative cases. This study aims to evaluate the effects of the Synthetic Minority Over-sampling Technique (SMOTE) and Ant Colony Optimization (ACO) on XGBoost performance and to investigate the trade-off between sensitivity, specificity, and overall discriminative ability. The Cervical Cancer (Risk Factors) dataset from the UCI Machine Learning Repository, containing 858 observations, was used with Biopsy as the binary classification target, comprising 803 negative and 55 positive cases. A total of 28 risk-factor features were retained after excluding diagnostic attributes with potential data leakage. Three classification scenarios were evaluated: XGBoost, SMOTE-XGBoost, and SMOTE-ACO-XGBoost. Model performance was assessed using Nested Repeated Stratified Cross-Validation, consisting of an outer 5-fold cross-validation repeated five times and an inner 3-fold cross-validation for hyperparameter and decision-threshold optimization. The results showed that SMOTE-XGBoost achieved the highest recall of 0.4182 ± 0.3340, whereas SMOTE-ACO-XGBoost obtained the highest accuracy of 0.7667 ± 0.1698 and specificity of 0.8042 ± 0.1950. Baseline XGBoost achieved the highest ROC-AUC and PR-AUC values of 0.6227 ± 0.0663 and 0.1290 ± 0.0739, respectively. The Friedman test indicated significant differences among the models across all evaluation metrics. Pairwise Wilcoxon signed-rank tests with Holm correction showed that ACO significantly improved accuracy and specificity compared with SMOTE-XGBoost but significantly reduced recall from 0.4182 to 0.2182. These findings demonstrate that class-imbalance handling and hyperparameter optimization do not produce a single model that consistently dominates all performance measures. SMOTE tends to improve sensitivity, whereas ACO optimization shifts the classifier toward higher accuracy and specificity at the cost of reduced sensitivity.

GS Cited Analysis

Downloads

Download data is not yet available.

References

Abdelhay, E. H., Elgamily, K. M., & Walaa Omar El-Farouk Badr. (2026). Metaheuristic optimization of deep CNNs for multi-class diagnosis of cervical cancer and lymphoma. Scientific Reports, 16. https://doi.org/https://doi.org/10.1038/s41598-026-51619-3

AlMohimeed, A., Saleh, H., Mostafa, S., Redhwan M. A. Saad, & Talaat, A. S. (2023). Cervical Cancer Diagnosis Using Stacked Ensemble Model and Optimized Feature Selection: An Explainable Artificial Intelligence Approach. Computers, 12(10). https://doi.org/https://doi.org/10.3390/computers12100200

Altalhan, M., Algarni, A., & Monia, T. H. A. (2025). Imbalanced Data Problem in Machine Learning: A Review. Turki Hadj Alouane Monia, 11. https://doi.org/10.1109/ACCESS.2025.3531662

Anggitasyah, D., & Siregar, M. A. P. (2023). Penerapan Metode Smote Extreme Gradient Boosting Untuk Klasifikasi Penyakit Kanker Serviks Di Kota Medan. Jutisi: Jurnal Ilmiah Teknik Informatika Dan Sistem Informasi, 12(2). https://doi.org/https://doi.org/10.35889/jutisi.v12i2.1479

Boldini, D., Grisoni, F., Kuhn, D., Friedrich, L., & Sieber, S. A. (2023). Practical guidelines for the use of gradient boosting for molecular property prediction. Journal of Cheminformatics, 15(1), 73. https://doi.org/10.1186/s13321-023-00743-7

Carolina, I., Andharsaputri, R. L., Suharjanti, S., Prihatin, T., & Nurdin, H. (2026). Comparative Analysis of Multi-Classifier Models with Resampling Techniques for Imbalanced Student Graduation Prediction. Paradigma, 28(1). https://doi.org/https://doi.org/10.31294/p.v28i1.11976

Çorbacıoğlu, Ş. K., & Aksel, G. (2023). Receiver operating characteristic curve analysis in diagnostic accuracy studies. Turkish Journal of Emergency Medicine, 23(4). https://doi.org/10.4103/tjem.tjem_182_23

Fernandes, K., Cardoso, J. S., & Fernandes, J. C. (2017). Cervical Cancer (Risk Factors). Springer. UCI Machine Learning Repository. https://doi.org/10.1007/978-3-319-58838-4_27

Fulazzaky, T., Saefuddin, A., & Soleh, A. M. (2024). Evaluating Ensemble Learning Techniques for Class Imbalance in Machine Learning: A Comparative Analysis of Balanced Random Forest, SMOTE-RF, SMOTEBoost, and RUSBoost. Scientific Journal of Informatics, 11(4). https://doi.org/https://doi.org/10.15294/sji.v11i4.15937

Geron, A. (2022). Hands on Machine learning with Scikit Learn Keras and Tensor Flow Concepts, Tools, and Techniques to Build Intelligent Systems (3rd ed.). O’Reilly Media, Inc.

Glučina, M., Ariana Lorencin, Nikola Anđelić, & Ivan Lorencin. (2023). Cervical Cancer Diagnostics Using Machine Learning Algorithms and Class Balancing Techniques. Applied Sciences, 13(12). https://doi.org/https://doi.org/10.3390/app13021061

Gurcan, F., & Soylu, A. (2024). Learning from Imbalanced Data: Integration of Advanced Resampling Techniques and Machine Learning Models for Enhanced Cancer Diagnosis and Prognosis. Cancers, 16(19). https://doi.org/https://doi.org/10.3390/cancers16193417

Hamid, A., & Ridwansyah. (2024). Optimizing Heart Failure Detection : A Comparison between Naive Bayes and Particle Swarm Optimization. Paradigma, 26(1), 30–36. https://doi.org/https://doi.org/10.31294/p.v26i1.3284

Huang, C. Y., & Dai, H. L. (2021). Learning from class-imbalanced data: review of data driven methods and algorithm driven methods. Data Science in Finance and Economics, 1(1), 21–36. https://doi.org/10.3934/DSFE.2021002

Karamti, H., Alharthi, R., Anizi, A. Al, Reemah M Alhebshi, Eshmawi, A. A., Alsubai, S., & Muhammad Umer. (2023). Improving Prediction of Cervical Cancer Using KNN Imputed SMOTE Features and Multi-Model Ensemble Learning Approach. Cancers, 15(17). https://doi.org/10.3390/cancers15174412

Kondo, T. S., Daniel Ngondya, & Hamim Rusheke. (2025). Application of artificial intelligence in cervical cancer diagnosis using risk factors: A systematic review. Telematics and Informatics Reports, 20. https://doi.org/https://doi.org/10.1016/j.teler.2025.100250

Marinho, T. L., Diego Carvalho do Nascimento, & Bruno Almeida Pimentel. (2024). Optimization on selecting XGBoost hyperparameters using meta-learning. Expert Systems, 41(9). https://doi.org/https://doi.org/10.1111/exsy.13611

Mudawi, N. Al, & Alazeb, A. (2022). A Model for Predicting Cervical Cancer Using Machine Learning Algorithms. Sensors, 22(11). https://doi.org/https://doi.org/10.3390/s22114132

Muraru, M. M., Simó, Z., & Iantovics, L. B. (2024). Cervical Cancer Prediction Based on Imbalanced Data Using Machine Learning Algorithms with a Variety of Sampling Methods. Applied Sciences, 14(22). https://doi.org/https://doi.org/10.3390/app142210085

Ridwansyah, R., Rahayu, S., Purnama, J. J., Riyanto, V., & Hamid, A. (2026). Pengoptimalan Seleksi Fitur Berbasis Particle Swarm Optimization pada Prediksi Gagal Jantung dengan Random Tree. JASIEK (Jurnal Aplikasi Sains, Informasi, Elektronika Dan Komputer), 8(1). https://doi.org/https://doi.org/10.26905/jasiek.v8i1.16595

Ridwansyah, R., Riyanto, V., Hamid, A., Rahayu, S., & Purnama, J. J. (2022). Grouping Data in Predicting Infant Mortality Using K-Means and Decision Tree. Paradigma, 24(2), 168–174. https://doi.org/10.31294/paradigma.v24i2.1399

Salmi, M., Atif, D., Oliva, D., Abraham, A., & Sebastian Ventura. (2024). Handling imbalanced medical datasets: review of a decade of research. Artificial Intelligence Review (Springer Nature, 57(10). https://doi.org/10.1007/s10462-024-10884-2

Saputra, R. M., Alzami, F., Pramudi, Y. T. C., Erawan, L., Megantara, R. A., Ricardus Anggi Pramunendar, & Yusuf, M. (2025). Improving Cervical Cancer Classification Using ADASYN and Random Forest with GridSearchCV Optimization. Informatics, Electrical Engineering, and Mechanical Engineering, 16(1). https://doi.org/https://doi.org/10.35970/infotekmesin.v16i1.2552

Sumarna, Astrilyana, Sugiono, Wijaya, G., & Yessica Fara Desvia. (2026). Improving Minority Class Detection in Cervical Cancer Prediction Using Imbalance-Aware Ensemble Learning. Sinkron : Jurnal Dan Penelitian Teknik Informatika, 10(2). https://doi.org/10.33395/sinkron.v10i2.15995

Sung, H., Filho, A. M., Laversanne, M., Ferlay, J., Rebecca L Siegel, Soerjomataram, I., Ahmedin Jemal, & Bray, F. (2026). Global cancer statistics 2024: GLOBOCAN estimates of incidence and mortality worldwide for 34 cancers in 186 countries. CA: A Cancer Journal for Clinicians, 76(4). https://doi.org/10.3322/caac.70090

Vazquez, B., Rojas-García, M., Rodríguez-Esquivel, J. I., Marquez-Acosta, J., Aranda-Flores, C. E., Cetina-Pérez, L. del C., Soto-López, S., Estévez-García, J. A., Bahena-Román, M., Madrid-Marina, V., & Torres-Poveda, K. (2025). Machine and Deep Learning for the Diagnosis, Prognosis, and Treatment of Cervical Cancer: A Scoping Review. Diagnostics, 15(12).

Yang, Y., Khorshidi, H. A., & Aickelin, U. (2024). A review on over-sampling techniques in classification of multi-class imbalanced datasets: insights for medical problems. Frontiers in Digital Health, 6(1430245). https://doi.org/10.3389/fdgth.2024.1430245

Yudha, M. A. R., & Rahardi, M. (2025). Comparative Analysis of Random Forest and XGBoost Models for Cervical Cancer Risk Prediction using SHAP-based Explainable AI. Journal of Applied Informatics and Computing, 9(6). https://doi.org/https://doi.org/10.30871/jaic.v9i6.10357

Downloads


Crossmark Updates

How to Cite

Prihatin, T., Suharjanti, S., Andharsaputri, R. L., & Nurdin, H. (2026). Evaluating the Sensitivity–Specificity Trade-off in Risk Factor-Based Cervical Cancer Classification Using SMOTE and ACO-Optimized XGBoost. Sinkron : Jurnal Dan Penelitian Teknik Informatika, 10(4), 2307-2318. https://doi.org/10.33395/sinkron.v10i4.16797