Application of Sentence-BERT Embeddings for Semantic Deduplication of Industrial Material Records
DOI:
10.33395/sinkron.v10i3.16220Keywords:
cosine similarity, entity resolution, material master data, record deduplication, semantic embedding, Sentence-BERTAbstract
Industrial material master data in Enterprise Resource Planning (ERP) and Enterprise Asset Management (EAM) systems accumulates duplicate records that distort inventory, procurement, and analytics. Traditional deduplication relies on string-similarity measures such as Levenshtein, Jaro–Winkler, and TF-IDF cosine, which can struggle on catalogs mixing Indonesian and English terminology—e.g. Valve versus Keran—and on paraphrastic variants with different word order or abbreviation style. This study formally specifies a semantic deduplication pipeline that encodes material descriptions as sentence embeddings using Sentence-BERT (SBERT) and compares them via cosine similarity, then diagnostically evaluates the extent to which SBERT improves over those baselines. Following Design Science Research, the pipeline specifies normalisation, encoding with a multilingual paraphrase-tuned SBERT variant, and pairwise comparison within candidate sets produced by hybrid blocking; the diagnostic evaluation reports scores on the raw descriptions to expose baseline behaviour before domain-specific harmonisation. A sample of 291,000 records from two Indonesian industrial power plants motivates the design. On a diagnostic set of 100 record pairs derived from existing engineer-annotated duplicate markers, Jaro–Winkler achieves F1 = 0.925 (precision 1.000, recall 0.860) and SBERT achieves F1 = 0.875 (precision 0.913, recall 0.840) at threshold τ = 0.65; qualitative analysis of twelve representative pairs further reveals that SBERT excels on structural paraphrase (cosine 0.73–0.88 where character-level methods score below 0.50), while Jaro–Winkler remains competitive on abbreviation, unit-standard, and cross-language pairs—particularly those involving Indonesian technical vocabulary under-represented in the model’s training distribution. The central finding is that Sentence-BERT complements rather than replaces string baselines, which motivates future work on multi-channel architectures combining textual semantics with structural context.
Downloads
References
Amin, M. M., Stiawan, D., Ermatita, & Budiarto, R. (2024). Komparasi kinerja algoritma blocking pada proses indexing untuk deteksi duplikasi. Jurnal Teknologi Informasi dan Ilmu Komputer, 11(4), 715–722. https://doi.org/10.25126/jtiik.1148080
Cahyawijaya, S., Lovenia, H., Aji, A. F., Winata, G., Wilie, B., Koto, F., … & Purwarianti, A. (2023). NusaCrowd: Open source initiative for Indonesian NLP resources. In Findings of the Association for Computational Linguistics: ACL 2023 (pp. 13745–13818). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.findings-acl.868
Chen, T. Y., Kuo, F.-C., Liu, H., Poon, P.-L., Towey, D., Tse, T. H., & Zhou, Z. Q. (2018). Metamorphic testing: A review of challenges and opportunities. ACM Computing Surveys, 51(1), 1–27. https://doi.org/10.1145/3143561
Christen, P. (2012). Data matching: Concepts and techniques for record linkage, entity resolution, and duplicate detection. Springer.
Christophides, V., Efthymiou, V., Palpanas, T., Papadakis, G., & Stefanidis, K. (2020). An overview of end-to-end entity resolution for big data. ACM Computing Surveys, 53(6), 1–42. https://doi.org/10.1145/3418896
Dell, M., & Therapontos, A. (2024). LinkTransformer: A unified package for record linkage with transformer language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics: System Demonstrations (ACL 2024) (pp. 222–233). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.acl-demos.21
Diana, D., & Khodra, M. L. (2023). IndoSBERT: Enhancing Indonesian sentence embeddings with Siamese networks fine-tuning. In Proceedings of the 10th International Conference on Advanced Informatics: Concept, Theory and Application (ICAICTA) (pp. 1–6). IEEE. https://doi.org/10.1109/ICAICTA59291.2023.10390469
Ebraheem, M., Thirumuruganathan, S., Joty, S., Ouzzani, M., & Tang, N. (2018). Distributed representations of tuples for entity resolution. Proceedings of the VLDB Endowment, 11(11), 1454–1467. https://doi.org/10.14778/3236187.3236198
Feng, F., Yang, Y., Cer, D., Arivazhagan, N., & Wang, W. (2022). Language-agnostic BERT sentence embedding. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL 2022) (pp. 878–891). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.acl-long.62
Hamza, A. R. (2024). Product matching using Sentence-BERT: A deep learning approach to e-commerce product deduplication. Engineering and Technology Journal, 9(12), 1–12. https://doi.org/10.5281/zenodo.14524722
He, S., Peng, N., Qiu, L., & Zhang, X. (2023). JointMatcher: Numerically-aware entity matching using pre-trained language models with attention concentration. Neurocomputing, 522, 207–219. https://doi.org/10.1016/j.neucom.2022.12.024
Hevner, A. R., March, S. T., Park, J., & Ram, S. (2004). Design science in information systems research. MIS Quarterly, 28(1), 75–105. https://doi.org/10.2307/25148625
Koto, F., Rahimi, A., Lau, J. H., & Baldwin, T. (2020). IndoLEM and IndoBERT: A benchmark dataset and pre-trained language model for Indonesian NLP. In Proceedings of the 28th International Conference on Computational Linguistics (COLING) (pp. 757–770). International Committee on Computational Linguistics. https://doi.org/10.18653/v1/2020.coling-main.66
Li, Y., Li, J., Suhara, Y., Doan, A., & Tan, W.-C. (2020). Deep entity matching with pre-trained language models. Proceedings of the VLDB Endowment, 14(1), 50–60. https://doi.org/10.14778/3421424.3421431
Mudgal, S., Li, H., Rekatsinas, T., Doan, A., Park, Y., Krishnan, G., Deep, R., Arcaute, E., & Raghavendra, V. (2018). Deep learning for entity matching: A design space exploration. In Proceedings of the 2018 International Conference on Management of Data (SIGMOD) (pp. 19–34). ACM. https://doi.org/10.1145/3183713.3196926
Papadakis, G., Skoutas, D., Thanos, E., & Palpanas, T. (2020). Blocking and filtering techniques for entity resolution: A survey. ACM Computing Surveys, 53(2), 31:1–31:42. https://doi.org/10.1145/3377455
Peeters, R., Steiner, A., & Bizer, C. (2025). Entity matching using large language models. In Proceedings of the 28th International Conference on Extending Database Technology (EDBT) (pp. 171–184). https://doi.org/10.48786/edbt.2025.15
Peffers, K., Tuunanen, T., Rothenberger, M. A., & Chatterjee, S. (2007). A design science research methodology for information systems research. Journal of Management Information Systems, 24(3), 45–77. https://doi.org/10.2753/MIS0742-1222240302
Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 3982–3992). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1410
Reimers, N., & Gurevych, I. (2020). Making monolingual sentence embeddings multilingual using knowledge distillation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 4512–4525). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.emnlp-main.365
Sun, Z., Hu, W., & Zhang, C. (2020). A benchmarking study of embedding-based entity alignment for knowledge graphs. Proceedings of the VLDB Endowment, 13(11), 2326–2340. https://doi.org/10.14778/3407790.3407828
Wang, Z., Sisman, B., Wei, H., & Dong, X. L. (2021). CorDEL: A contrastive deep learning approach for entity linkage. In Proceedings of the 2021 IEEE International Conference on Data Engineering (ICDE) (pp. 1027–1038). IEEE. https://doi.org/10.1109/ICDE51399.2021.00093
Wilie, B., Vincentio, K., Winata, G. I., Cahyawijaya, S., Li, X., Lim, Z. Y., … & Purwarianti, A. (2020). IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing (AACL-IJCNLP) (pp. 843–857). Association for Computational Linguistics.
Zeakis, A., Papadakis, G., Skoutas, D., & Koubarakis, M. (2025). An in-depth analysis of pre-trained embeddings for entity resolution. The VLDB Journal, 34(1), 5. https://doi.org/10.1007/s00778-024-00879-4
Downloads
How to Cite
Issue
Section
License
Copyright (c) 2026 Seno Hardijanto Purnomo, Agung Triayudi

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.






















Moraref
PKP Index
Indonesia OneSearch
OCLC Worldcat
Index Copernicus
Scilit
