Indonesian Hate Speech Detection Across Diverse Domains Using Parameter-Efficient Fine-Tuning with IndoBERT and LoRA

Authors

  • Fergie Joanda Kaunang Universitas Bunda Mulia
  • Bhustomy Hakim Universitas Bunda Mulia
  • Angelina Pramana Thenata Universitas Bunda Mulia

DOI:

10.33395/jmp.v15i2.16581

Keywords:

Back Translation, IndoBERT, Indonesian Hate Speech, Low-Rank Adaptation (LoRA), Multi-domain Classification

Abstract

The rapid proliferation of digital connectivity in Indonesia has catalyzed an unprecedented surge in harmful online content, necessitating robust automated systems for hate speech detection that can generalize across diverse digital platforms. Traditional models often struggle with domain shift and the linguistic complexities of Indonesian social media discourse, including informal slang and code-mixing. This research proposes a multi-domain detection framework leveraging the IndoBERT-base-p1 architecture integrated with Low-Rank Adaptation (LoRA), a parameter-efficient fine-tuning (PEFT) strategy. The study utilizes a multi-source corpus from Instagram, Twitter, and news portals, employing back-translation to augment scarce Instagram data and stratified downsampling to ensure domain equilibrium. By training only 1–2% of the total 110 million parameters, specifically targeting the query and value attention modules, the model achieves significant computational savings with a training loss of 0.345. Experimental results demonstrate high robustness, with the framework attaining F1-scores of 0.83 for both Instagram and Twitter, and 0.81 for news portals, while maintaining accuracies between 0.81 and 0.85. Qualitative validation through word cloud analysis further confirms the model's ability to distinguish between aggressive sociopolitical triggers and neutral functional discourse. This study contributes a scalable and resource-efficient solution for real-time content moderation, proving effective across both formal journalistic Indonesian and informal digital dialects. The findings indicate that IndoBERT+LoRA provides a promising and resource-efficient approach for multi-domain Indonesian hate and abusive speech classification, while stricter leave-one-domain-out evaluation remains an important direction for future work.

GS Cited Analysis

Downloads

How to Cite

Kaunang, F. J., Hakim, B. ., & Thenata, A. P. . (2026). Indonesian Hate Speech Detection Across Diverse Domains Using Parameter-Efficient Fine-Tuning with IndoBERT and LoRA. Jurnal Minfo Polgan, 15(2), 2369-2378. https://doi.org/10.33395/jmp.v15i2.16581