Application of Sentence-BERT Embeddings for Semantic Deduplication of Industrial Material Records

Authors

  • Seno Hardijanto Purnomo Universitas Nasional
  • Agung Triayudi Universitas Nasional

DOI:

10.33395/sinkron.v10i3.16220

Keywords:

cosine similarity, entity resolution, material master data, record deduplication, semantic embedding, Sentence-BERT

Abstract

Industrial material master data in Enterprise Resource Planning (ERP) and Enterprise Asset Management (EAM) systems accumulates duplicate records that distort inventory, procurement, and analytics. Traditional deduplication relies on string-similarity measures such as Levenshtein, Jaro–Winkler, and TF-IDF cosine, which can struggle on catalogs mixing Indonesian and English terminology—e.g. Valve versus Keran—and on paraphrastic variants with different word order or abbreviation style. This study formally specifies a semantic deduplication pipeline that encodes material descriptions as sentence embeddings using Sentence-BERT (SBERT) and compares them via cosine similarity, then diagnostically evaluates the extent to which SBERT improves over those baselines. Following Design Science Research, the pipeline specifies normalisation, encoding with a multilingual paraphrase-tuned SBERT variant, and pairwise comparison within candidate sets produced by hybrid blocking; the diagnostic evaluation reports scores on the raw descriptions to expose baseline behaviour before domain-specific harmonisation. A sample of 291,000 records from two Indonesian industrial power plants motivates the design. On a diagnostic set of 100 record pairs derived from existing engineer-annotated duplicate markers, Jaro–Winkler achieves F1 = 0.925 (precision 1.000, recall 0.860) and SBERT achieves F1 = 0.875 (precision 0.913, recall 0.840) at threshold τ = 0.65; qualitative analysis of twelve representative pairs further reveals that SBERT excels on structural paraphrase (cosine 0.73–0.88 where character-level methods score below 0.50), while Jaro–Winkler remains competitive on abbreviation, unit-standard, and cross-language pairs—particularly those involving Indonesian technical vocabulary under-represented in the model’s training distribution. The central finding is that Sentence-BERT complements rather than replaces string baselines, which motivates future work on multi-channel architectures combining textual semantics with structural context.

 

GS Cited Analysis

Downloads

Download data is not yet available.

References

Amin, M. M., Stiawan, D., Ermatita, & Budiarto, R. (2024). Komparasi kinerja algoritma blocking pada proses indexing untuk deteksi duplikasi. Jurnal Teknologi Informasi dan Ilmu Komputer, 11(4), 715–722. https://doi.org/10.25126/jtiik.1148080

Cahyawijaya, S., Lovenia, H., Aji, A. F., Winata, G., Wilie, B., Koto, F., … & Purwarianti, A. (2023). NusaCrowd: Open source initiative for Indonesian NLP resources. In Findings of the Association for Computational Linguistics: ACL 2023 (pp. 13745–13818). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.findings-acl.868

Chen, T. Y., Kuo, F.-C., Liu, H., Poon, P.-L., Towey, D., Tse, T. H., & Zhou, Z. Q. (2018). Metamorphic testing: A review of challenges and opportunities. ACM Computing Surveys, 51(1), 1–27. https://doi.org/10.1145/3143561

Christen, P. (2012). Data matching: Concepts and techniques for record linkage, entity resolution, and duplicate detection. Springer.

Christophides, V., Efthymiou, V., Palpanas, T., Papadakis, G., & Stefanidis, K. (2020). An overview of end-to-end entity resolution for big data. ACM Computing Surveys, 53(6), 1–42. https://doi.org/10.1145/3418896

Dell, M., & Therapontos, A. (2024). LinkTransformer: A unified package for record linkage with transformer language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics: System Demonstrations (ACL 2024) (pp. 222–233). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.acl-demos.21

Diana, D., & Khodra, M. L. (2023). IndoSBERT: Enhancing Indonesian sentence embeddings with Siamese networks fine-tuning. In Proceedings of the 10th International Conference on Advanced Informatics: Concept, Theory and Application (ICAICTA) (pp. 1–6). IEEE. https://doi.org/10.1109/ICAICTA59291.2023.10390469

Ebraheem, M., Thirumuruganathan, S., Joty, S., Ouzzani, M., & Tang, N. (2018). Distributed representations of tuples for entity resolution. Proceedings of the VLDB Endowment, 11(11), 1454–1467. https://doi.org/10.14778/3236187.3236198

Feng, F., Yang, Y., Cer, D., Arivazhagan, N., & Wang, W. (2022). Language-agnostic BERT sentence embedding. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL 2022) (pp. 878–891). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.acl-long.62

Hamza, A. R. (2024). Product matching using Sentence-BERT: A deep learning approach to e-commerce product deduplication. Engineering and Technology Journal, 9(12), 1–12. https://doi.org/10.5281/zenodo.14524722

He, S., Peng, N., Qiu, L., & Zhang, X. (2023). JointMatcher: Numerically-aware entity matching using pre-trained language models with attention concentration. Neurocomputing, 522, 207–219. https://doi.org/10.1016/j.neucom.2022.12.024

Hevner, A. R., March, S. T., Park, J., & Ram, S. (2004). Design science in information systems research. MIS Quarterly, 28(1), 75–105. https://doi.org/10.2307/25148625

Koto, F., Rahimi, A., Lau, J. H., & Baldwin, T. (2020). IndoLEM and IndoBERT: A benchmark dataset and pre-trained language model for Indonesian NLP. In Proceedings of the 28th International Conference on Computational Linguistics (COLING) (pp. 757–770). International Committee on Computational Linguistics. https://doi.org/10.18653/v1/2020.coling-main.66

Li, Y., Li, J., Suhara, Y., Doan, A., & Tan, W.-C. (2020). Deep entity matching with pre-trained language models. Proceedings of the VLDB Endowment, 14(1), 50–60. https://doi.org/10.14778/3421424.3421431

Mudgal, S., Li, H., Rekatsinas, T., Doan, A., Park, Y., Krishnan, G., Deep, R., Arcaute, E., & Raghavendra, V. (2018). Deep learning for entity matching: A design space exploration. In Proceedings of the 2018 International Conference on Management of Data (SIGMOD) (pp. 19–34). ACM. https://doi.org/10.1145/3183713.3196926

Papadakis, G., Skoutas, D., Thanos, E., & Palpanas, T. (2020). Blocking and filtering techniques for entity resolution: A survey. ACM Computing Surveys, 53(2), 31:1–31:42. https://doi.org/10.1145/3377455

Peeters, R., Steiner, A., & Bizer, C. (2025). Entity matching using large language models. In Proceedings of the 28th International Conference on Extending Database Technology (EDBT) (pp. 171–184). https://doi.org/10.48786/edbt.2025.15

Peffers, K., Tuunanen, T., Rothenberger, M. A., & Chatterjee, S. (2007). A design science research methodology for information systems research. Journal of Management Information Systems, 24(3), 45–77. https://doi.org/10.2753/MIS0742-1222240302

Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 3982–3992). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1410

Reimers, N., & Gurevych, I. (2020). Making monolingual sentence embeddings multilingual using knowledge distillation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 4512–4525). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.emnlp-main.365

Sun, Z., Hu, W., & Zhang, C. (2020). A benchmarking study of embedding-based entity alignment for knowledge graphs. Proceedings of the VLDB Endowment, 13(11), 2326–2340. https://doi.org/10.14778/3407790.3407828

Wang, Z., Sisman, B., Wei, H., & Dong, X. L. (2021). CorDEL: A contrastive deep learning approach for entity linkage. In Proceedings of the 2021 IEEE International Conference on Data Engineering (ICDE) (pp. 1027–1038). IEEE. https://doi.org/10.1109/ICDE51399.2021.00093

Wilie, B., Vincentio, K., Winata, G. I., Cahyawijaya, S., Li, X., Lim, Z. Y., … & Purwarianti, A. (2020). IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing (AACL-IJCNLP) (pp. 843–857). Association for Computational Linguistics.

Zeakis, A., Papadakis, G., Skoutas, D., & Koubarakis, M. (2025). An in-depth analysis of pre-trained embeddings for entity resolution. The VLDB Journal, 34(1), 5. https://doi.org/10.1007/s00778-024-00879-4

Downloads


Crossmark Updates

How to Cite

Purnomo, S. H., & Triayudi, A. . (2026). Application of Sentence-BERT Embeddings for Semantic Deduplication of Industrial Material Records. Sinkron : Jurnal Dan Penelitian Teknik Informatika, 10(3), 1404-1411. https://doi.org/10.33395/sinkron.v10i3.16220