Bayesian-Optimized Word2Vec-BiLSTM for Malicious Prompt Detection in LLMs

Authors

  • Hilman Singgih Wicaksana Informatics Program, Universitas Karya Husada, Semarang, Indonesia
  • Gregorius Airlangga Information Systems Program, Universitas Katolik Indonesia Atma Jaya, Jakarta, Indonesia

DOI:

10.33395/sinkron.v10i4.16700

Keywords:

Bayesian Optimization, BiLSTM, FastText, LLM Security, Prompt Injection Detection

Abstract

Large Language Models (LLMs) are increasingly adopted in various artificial intelligence applications, but their growing use also raises concerns about LLM security, particularly against prompt injection attacks that manipulate input instructions to influence model behavior. This study aims to develop an efficient and effective prompt injection detection model by proposing FastText-BiLSTM, a lightweight approach for classifying prompts as malicious or non-malicious. The proposed model combines FastText embeddings, which capture subword-level information, with a BiLSTM that learns sequential patterns in both directions. To provide a comprehensive evaluation, this study also compares the proposed model with three alternative combinations, namely Word2Vec-BiLSTM, FastText-BiGRU, and Word2Vec-BiGRU. Bayesian Optimization was applied to obtain the optimal hyperparameter configuration for each model, and performance was evaluated using precision, recall, F1-score, accuracy, and confusion matrix analysis. The results show that FastText-BiLSTM achieved the best performance, with a precision of 96.77% and recall, F1-score, and accuracy of approximately 96.76%. Compared with other models, FastText-BiLSTM demonstrated a more balanced classification capability and fewer errors when distinguishing malicious from non-malicious prompts. These findings indicate that the combination of FastText and BiLSTM provides competitive detection performance while maintaining a relatively lightweight architecture. Therefore, FastText-BiLSTM can be considered the most accurate and optimal model in this study, demonstrating its effectiveness as a lightweight deep learning approach for supporting LLM security against prompt injection attacks.

GS Cited Analysis

Downloads

Download data is not yet available.

References

Ahmed, S. F., Alam, Md. S. Bin, Hassan, M., Rozbu, M. R., Ishtiak, T., Rafa, N., Mofijur, M., Shawkat Ali, A. B. M., & Gandomi, A. H. (2023). Deep learning modelling techniques: current progress, applications, advantages, and challenges. Artificial Intelligence Review, 56(11), 13521–13617. https://doi.org/10.1007/s10462-023-10466-8

Ali, M. D., Saleem, A., Elahi, H., Khan, M. A., Khan, M. I., Yaqoob, M. M., Farooq Khattak, U., & Al-Rasheed, A. (2023). Breast Cancer Classification through Meta-Learning Ensemble Technique Using Convolution Neural Networks. Diagnostics, 13(13), 2242. https://doi.org/10.3390/diagnostics13132242

Alshammari, A. A., & Alsaleh, O. I. (2026). Detecting Prompt Injection Attacks in Generative AI Systems: A Hybrid SIEM and One-Class SVM Framework. Electronics, 15(11), 2242. https://doi.org/10.3390/electronics15112242

Bokolo, B. G., & Liu, Q. (2023). Deep Learning-Based Depression Detection from Social Media: Comparative Evaluation of ML and Transformer Techniques. Electronics, 12(21), 4396. https://doi.org/10.3390/electronics12214396

Bolikulov, F., Nasimov, R., Rashidov, A., Akhmedov, F., & Cho, Y.-I. (2024). Effective Methods of Categorical Data Encoding for Artificial Intelligence Algorithms. Mathematics, 12(16), 2553. https://doi.org/10.3390/math12162553

Bongirwar, V., & Mokhade, A. S. (2024). A Hybrid Bidirectional Long Short-Term Memory and Bidirectional Gated Recurrent Unit Architecture for Protein Secondary Structure Prediction. IEEE Access, 12, 115346–115355. https://doi.org/10.1109/ACCESS.2024.3444468

Cihan, P. (2025). Bayesian Hyperparameter Optimization of Machine Learning Models for Predicting Biomass Gasification Gases. Applied Sciences, 15(3), 1018. https://doi.org/10.3390/app15031018

Derner, E., Batistič, K., Zahálka, J., & Babuška, R. (2024). A Security Risk Taxonomy for Prompt-Based Interaction With Large Language Models. IEEE Access, 12, 126176–126187. https://doi.org/10.1109/ACCESS.2024.3450388

Feretzakis, G., & Verykios, V. S. (2024). Trustworthy AI: Securing Sensitive Data in Large Language Models. AI, 5(4), 2773–2800. https://doi.org/10.3390/ai5040134

Jebbar, M. A. (2025). Malicious Prompt Detection Dataset (MPDD).

Jelodar, M. B. (2025). Generative AI, Large Language Models, and ChatGPT in Construction Education, Training, and Practice. Buildings, 15(6), 933. https://doi.org/10.3390/buildings15060933

Kurniawan, A., & Chandra, M. B. (2025). Simulation of prompt injection attacks on generative pre-trained transformers models. Procedia Computer Science, 269, 400–410. https://doi.org/10.1016/j.procs.2025.08.292

Kushnerov, O., Shevchuk, R., Yevseiev, S., & Karpiński, M. (2026). Comparative Benchmarking of Deep Learning Architectures for Detecting Adversarial Attacks on Large Language Models. Information (Switzerland), 17(2). https://doi.org/10.3390/info17020155

Lan, Q., Kaul, A., & Jones, S. (2025). Prompt Injection Detection in LLM Integrated Applications. International Journal of Network Dynamics and Intelligence, 4(2). https://doi.org/10.53941/ijndi.2025.100013

Li, L., Li, J., Wang, H., & Nie, J. (2024). Application of the transformer model algorithm in chinese word sense disambiguation: a case study in chinese language. Scientific Reports, 14(1), 6320. https://doi.org/10.1038/s41598-024-56976-5

Lubis, A. R., Lase, Y. Y., Rahman, D. A., & Witarsyah, D. (2023). Improving Spell Checker Performance for Bahasa Indonesia Using Text Preprocessing Techniques with Deep Learning Models. Ingénierie Des Systèmes d Information, 28(5), 1335–1342. https://doi.org/10.18280/isi.280522

Mirshekali, H., Reza Shadi, M., Ghanadi Ladani, F., & Reza Shaker, H. (2025). A Review of Large Language Models for Energy Systems: Applications, Challenges, and Future Prospects. IEEE Access, 13, 163162–163188. https://doi.org/10.1109/ACCESS.2025.3610994

Mutinda, J., Mwangi, W., & Okeyo, G. (2023). Sentiment Analysis of Text Reviews Using Lexicon-Enhanced Bert Embedding (LeBERT) Model with Convolutional Neural Network. Applied Sciences, 13(3), 1445. https://doi.org/10.3390/app13031445

Patil, R., Boit, S., Gudivada, V., & Nandigam, J. (2023). A Survey of Text Representation and Embedding Techniques in NLP. IEEE Access, 11, 36120–36146. https://doi.org/10.1109/ACCESS.2023.3266377

Pavlatos, C., Makris, E., Fotis, G., Vita, V., & Mladenov, V. (2023). Enhancing Electrical Load Prediction Using a Bidirectional LSTM Neural Network. Electronics (Switzerland), 12(22). https://doi.org/10.3390/electronics12224652

Raza, M., Jahangir, Z., Riaz, M. B., Saeed, M. J., & Sattar, M. A. (2025). Industrial applications of large language models. Scientific Reports, 15(1), 13755. https://doi.org/10.1038/s41598-025-98483-1

Riyanto, S., Sitanggang, I. S., Djatna, T., & Atikah, T. D. (2023). Comparative Analysis using Various Performance Metrics in Imbalanced Data for Multi-class Text Classification. International Journal of Advanced Computer Science and Applications, 14(6). https://doi.org/10.14569/IJACSA.2023.01406116

Salehin, I., & Kang, D.-K. (2023). A Review on Dropout Regularization Approaches for Deep Neural Networks within the Scholarly Domain. Electronics, 12(14), 3106. https://doi.org/10.3390/electronics12143106

Shaday, E. N., Engel, V. J. L., & Heryanto, H. (2024). Application of the Bidirectional Long Short-Term Memory Method with Comparison of Word2Vec, GloVe, and FastText for Emotion Classification in Song Lyrics. Procedia Computer Science, 245, 137–146. https://doi.org/10.1016/j.procs.2024.10.237

Sujon, K. M., Hassan, R., Choi, K., & Samad, M. A. (2025). Accuracy, precision, recall, f1-score, or MCC? empirical evidence from advanced statistics, ML, and XAI for evaluating business predictive models. Journal of Big Data, 12(1), 268. https://doi.org/10.1186/s40537-025-01313-4

Varoquaux, G., & Colliot, O. (2023). Evaluating Machine Learning Models and Their Diagnostic Value. In O. Colliot (Ed.), Machine Learning for Brain Disorders (pp. 601–630). Springer US. https://doi.org/10.1007/978-1-0716-3195-9_20

Venkatesh, S., Sindhu, R., & Arunachalam, V. (2025). Hardware efficient approximate sigmoid activation function for classifying features around zero. Integration, 103, 102421. https://doi.org/10.1016/j.vlsi.2025.102421

Vyas, K. P., Bhatt, M., Dudhatra, D., Gupta, R., & Sharma, A. (2026). Malicious Prompt Classifier with Leave-One-Out Deletion Approach for Prompt Sanitization. Eighth International Conference on Futuristic Trends in Networks and Computing Technologies (FTNCT08), 3015–3024. www.sciencedirect.com

Wicaksana, H. S., Kusumaningrum, R., & Gernowo, R. (2024). Determining community happiness index with transformers and attention-based deep learning. IAES International Journal of Artificial Intelligence (IJ-AI), 13(2), 1753. https://doi.org/10.11591/ijai.v13.i2.pp1753-1761

Wicaksono, G. W., Al asqalani, S. F., Azhar, Y., Hidayah, N. P., & Andreawana, A. (2023). Automatic Summarization of Court Decision Documents over Narcotic Cases Using BERT. JOIV : International Journal on Informatics Visualization, 7(2), 416. https://doi.org/10.30630/joiv.7.2.1811

Downloads


Crossmark Updates

How to Cite

Wicaksana, H. S., & Airlangga, G. (2026). Bayesian-Optimized Word2Vec-BiLSTM for Malicious Prompt Detection in LLMs. Sinkron : Jurnal Dan Penelitian Teknik Informatika, 10(4), 2070-2086. https://doi.org/10.33395/sinkron.v10i4.16700