Bayesian-Optimized Word2Vec-BiLSTM for Malicious Prompt Detection in LLMs
DOI:
10.33395/sinkron.v10i4.16700Keywords:
Bayesian Optimization, BiLSTM, FastText, LLM Security, Prompt Injection DetectionAbstract
Large Language Models (LLMs) are increasingly adopted in various artificial intelligence applications, but their growing use also raises concerns about LLM security, particularly against prompt injection attacks that manipulate input instructions to influence model behavior. This study aims to develop an efficient and effective prompt injection detection model by proposing FastText-BiLSTM, a lightweight approach for classifying prompts as malicious or non-malicious. The proposed model combines FastText embeddings, which capture subword-level information, with a BiLSTM that learns sequential patterns in both directions. To provide a comprehensive evaluation, this study also compares the proposed model with three alternative combinations, namely Word2Vec-BiLSTM, FastText-BiGRU, and Word2Vec-BiGRU. Bayesian Optimization was applied to obtain the optimal hyperparameter configuration for each model, and performance was evaluated using precision, recall, F1-score, accuracy, and confusion matrix analysis. The results show that FastText-BiLSTM achieved the best performance, with a precision of 96.77% and recall, F1-score, and accuracy of approximately 96.76%. Compared with other models, FastText-BiLSTM demonstrated a more balanced classification capability and fewer errors when distinguishing malicious from non-malicious prompts. These findings indicate that the combination of FastText and BiLSTM provides competitive detection performance while maintaining a relatively lightweight architecture. Therefore, FastText-BiLSTM can be considered the most accurate and optimal model in this study, demonstrating its effectiveness as a lightweight deep learning approach for supporting LLM security against prompt injection attacks.
Downloads
References
Ahmed, S. F., Alam, Md. S. Bin, Hassan, M., Rozbu, M. R., Ishtiak, T., Rafa, N., Mofijur, M., Shawkat Ali, A. B. M., & Gandomi, A. H. (2023). Deep learning modelling techniques: current progress, applications, advantages, and challenges. Artificial Intelligence Review, 56(11), 13521–13617. https://doi.org/10.1007/s10462-023-10466-8
Ali, M. D., Saleem, A., Elahi, H., Khan, M. A., Khan, M. I., Yaqoob, M. M., Farooq Khattak, U., & Al-Rasheed, A. (2023). Breast Cancer Classification through Meta-Learning Ensemble Technique Using Convolution Neural Networks. Diagnostics, 13(13), 2242. https://doi.org/10.3390/diagnostics13132242
Alshammari, A. A., & Alsaleh, O. I. (2026). Detecting Prompt Injection Attacks in Generative AI Systems: A Hybrid SIEM and One-Class SVM Framework. Electronics, 15(11), 2242. https://doi.org/10.3390/electronics15112242
Bokolo, B. G., & Liu, Q. (2023). Deep Learning-Based Depression Detection from Social Media: Comparative Evaluation of ML and Transformer Techniques. Electronics, 12(21), 4396. https://doi.org/10.3390/electronics12214396
Bolikulov, F., Nasimov, R., Rashidov, A., Akhmedov, F., & Cho, Y.-I. (2024). Effective Methods of Categorical Data Encoding for Artificial Intelligence Algorithms. Mathematics, 12(16), 2553. https://doi.org/10.3390/math12162553
Bongirwar, V., & Mokhade, A. S. (2024). A Hybrid Bidirectional Long Short-Term Memory and Bidirectional Gated Recurrent Unit Architecture for Protein Secondary Structure Prediction. IEEE Access, 12, 115346–115355. https://doi.org/10.1109/ACCESS.2024.3444468
Cihan, P. (2025). Bayesian Hyperparameter Optimization of Machine Learning Models for Predicting Biomass Gasification Gases. Applied Sciences, 15(3), 1018. https://doi.org/10.3390/app15031018
Derner, E., Batistič, K., Zahálka, J., & Babuška, R. (2024). A Security Risk Taxonomy for Prompt-Based Interaction With Large Language Models. IEEE Access, 12, 126176–126187. https://doi.org/10.1109/ACCESS.2024.3450388
Feretzakis, G., & Verykios, V. S. (2024). Trustworthy AI: Securing Sensitive Data in Large Language Models. AI, 5(4), 2773–2800. https://doi.org/10.3390/ai5040134
Jebbar, M. A. (2025). Malicious Prompt Detection Dataset (MPDD).
Jelodar, M. B. (2025). Generative AI, Large Language Models, and ChatGPT in Construction Education, Training, and Practice. Buildings, 15(6), 933. https://doi.org/10.3390/buildings15060933
Kurniawan, A., & Chandra, M. B. (2025). Simulation of prompt injection attacks on generative pre-trained transformers models. Procedia Computer Science, 269, 400–410. https://doi.org/10.1016/j.procs.2025.08.292
Kushnerov, O., Shevchuk, R., Yevseiev, S., & Karpiński, M. (2026). Comparative Benchmarking of Deep Learning Architectures for Detecting Adversarial Attacks on Large Language Models. Information (Switzerland), 17(2). https://doi.org/10.3390/info17020155
Lan, Q., Kaul, A., & Jones, S. (2025). Prompt Injection Detection in LLM Integrated Applications. International Journal of Network Dynamics and Intelligence, 4(2). https://doi.org/10.53941/ijndi.2025.100013
Li, L., Li, J., Wang, H., & Nie, J. (2024). Application of the transformer model algorithm in chinese word sense disambiguation: a case study in chinese language. Scientific Reports, 14(1), 6320. https://doi.org/10.1038/s41598-024-56976-5
Lubis, A. R., Lase, Y. Y., Rahman, D. A., & Witarsyah, D. (2023). Improving Spell Checker Performance for Bahasa Indonesia Using Text Preprocessing Techniques with Deep Learning Models. Ingénierie Des Systèmes d Information, 28(5), 1335–1342. https://doi.org/10.18280/isi.280522
Mirshekali, H., Reza Shadi, M., Ghanadi Ladani, F., & Reza Shaker, H. (2025). A Review of Large Language Models for Energy Systems: Applications, Challenges, and Future Prospects. IEEE Access, 13, 163162–163188. https://doi.org/10.1109/ACCESS.2025.3610994
Mutinda, J., Mwangi, W., & Okeyo, G. (2023). Sentiment Analysis of Text Reviews Using Lexicon-Enhanced Bert Embedding (LeBERT) Model with Convolutional Neural Network. Applied Sciences, 13(3), 1445. https://doi.org/10.3390/app13031445
Patil, R., Boit, S., Gudivada, V., & Nandigam, J. (2023). A Survey of Text Representation and Embedding Techniques in NLP. IEEE Access, 11, 36120–36146. https://doi.org/10.1109/ACCESS.2023.3266377
Pavlatos, C., Makris, E., Fotis, G., Vita, V., & Mladenov, V. (2023). Enhancing Electrical Load Prediction Using a Bidirectional LSTM Neural Network. Electronics (Switzerland), 12(22). https://doi.org/10.3390/electronics12224652
Raza, M., Jahangir, Z., Riaz, M. B., Saeed, M. J., & Sattar, M. A. (2025). Industrial applications of large language models. Scientific Reports, 15(1), 13755. https://doi.org/10.1038/s41598-025-98483-1
Riyanto, S., Sitanggang, I. S., Djatna, T., & Atikah, T. D. (2023). Comparative Analysis using Various Performance Metrics in Imbalanced Data for Multi-class Text Classification. International Journal of Advanced Computer Science and Applications, 14(6). https://doi.org/10.14569/IJACSA.2023.01406116
Salehin, I., & Kang, D.-K. (2023). A Review on Dropout Regularization Approaches for Deep Neural Networks within the Scholarly Domain. Electronics, 12(14), 3106. https://doi.org/10.3390/electronics12143106
Shaday, E. N., Engel, V. J. L., & Heryanto, H. (2024). Application of the Bidirectional Long Short-Term Memory Method with Comparison of Word2Vec, GloVe, and FastText for Emotion Classification in Song Lyrics. Procedia Computer Science, 245, 137–146. https://doi.org/10.1016/j.procs.2024.10.237
Sujon, K. M., Hassan, R., Choi, K., & Samad, M. A. (2025). Accuracy, precision, recall, f1-score, or MCC? empirical evidence from advanced statistics, ML, and XAI for evaluating business predictive models. Journal of Big Data, 12(1), 268. https://doi.org/10.1186/s40537-025-01313-4
Varoquaux, G., & Colliot, O. (2023). Evaluating Machine Learning Models and Their Diagnostic Value. In O. Colliot (Ed.), Machine Learning for Brain Disorders (pp. 601–630). Springer US. https://doi.org/10.1007/978-1-0716-3195-9_20
Venkatesh, S., Sindhu, R., & Arunachalam, V. (2025). Hardware efficient approximate sigmoid activation function for classifying features around zero. Integration, 103, 102421. https://doi.org/10.1016/j.vlsi.2025.102421
Vyas, K. P., Bhatt, M., Dudhatra, D., Gupta, R., & Sharma, A. (2026). Malicious Prompt Classifier with Leave-One-Out Deletion Approach for Prompt Sanitization. Eighth International Conference on Futuristic Trends in Networks and Computing Technologies (FTNCT08), 3015–3024. www.sciencedirect.com
Wicaksana, H. S., Kusumaningrum, R., & Gernowo, R. (2024). Determining community happiness index with transformers and attention-based deep learning. IAES International Journal of Artificial Intelligence (IJ-AI), 13(2), 1753. https://doi.org/10.11591/ijai.v13.i2.pp1753-1761
Wicaksono, G. W., Al asqalani, S. F., Azhar, Y., Hidayah, N. P., & Andreawana, A. (2023). Automatic Summarization of Court Decision Documents over Narcotic Cases Using BERT. JOIV : International Journal on Informatics Visualization, 7(2), 416. https://doi.org/10.30630/joiv.7.2.1811
Downloads
How to Cite
Issue
Section
License
Copyright (c) 2026 Hilman Singgih Wicaksana, Gregorius Airlangga

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.






















Moraref
PKP Index
Indonesia OneSearch
OCLC Worldcat
Index Copernicus
Scilit
