Optimizing Scientific Document Similarity Detection Using Jaro–Winkler and Machine Learning
Abstract
Keywords
Full Text:
PDFReferences
Ahmad, T., Ahamad, M., Ahmed, S. U., Ahmad, N., & Ahmad, N. (2022). Short question-answers assessment using lexical and semantic similarity-based features. Journal of Discrete Mathematical Sciences and Cryptography, 25(7), 2057–2067.
Alian, M., & Awajan, A. (2021). Arabic sentence similarity based on similarity features and machine learning. Soft Computing, 25(15), 10089–10101.
Baloi, A., Belean, B., Turcu, F., & Peptenatu, D. (2024). GPU-based similarity metrics computation and machine learning approaches for string similarity evaluation in large datasets. Soft Computing, 28(4), 3465–3477.
Basile, A., Crupi, R., Grasso, M., Mercanti, A., Regoli, D., Scarsi, S., et al. (2024). Disambiguation of company names via deep recurrent networks. Expert Systems with Applications, 238, 122035.
Bird, S., Klein, E., & Loper, E. (2009). Natural language processing with Python. O'Reilly Media.
Christen, P. (2012). Data matching: Concepts and techniques for record linkage, entity resolution, and duplicate detection. Springer.
Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT 2019 (pp. 4171–4186). Association for Computational Linguistics.
Draschner, C. F., Jabeen, H., & Lehmann, J. (2022). SimE4KG: Distributed and explainable multi-modal semantic similarity estimation for knowledge graphs. In 2022 IEEE Fifth International Conference on Artificial Intelligence and Knowledge Engineering (AIKE) (pp. 1–8). IEEE.
Foltýnek, T., Meuschke, N., & Gipp, B. (2020). Academic plagiarism detection: A systematic literature review. ACM Computing Surveys, 52(6), 1–42.
Jaro, M. A. (1989). Advances in record-linkage methodology as applied to matching the 1985 census of Tampa, Florida. Journal of the American Statistical Association, 84(406), 414–420.
Kazemian, H., & Shrestha, S. (2023). Comparisons of machine learning techniques for detecting fraudulent criminal identities. Expert Systems with Applications, 229, 120591.
Kowsari, K., Jafari Meimandi, K., Heidarysafa, M., Mendu, S., Barnes, L., & Brown, D. (2019). Text classification algorithms: A survey. Information, 10(4), 150.
Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to information retrieval. Cambridge University Press.
Molloy, C., Banks, J., Ding, S. H. H., Alaca, F., Charland, P., & Walenstein, A. (2025). Mecha: A neural-symbolic open-set homogeneous decision fusion approach for zero-day malware similarity detection. IEEE Transactions on Software Engineering, 51(2), 621–637.
Novyantika, R. D., & Isa, S. M. (2023). Improve data text quality by applying text pre-processing method (case study). International Journal of Engineering Trends and Technology, 71(1), 94–108.
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830.
Pitchandi, P., & Balakrishnan, M. (2023). Document clustering analysis with aid of adaptive Jaro Winkler with Jellyfish search clustering algorithm. Advances in Engineering Software, 175, 103322.
Potthast, M., Hagen, M., Göring, S., Rosso, P., & Stein, B. (2014). Overview of the 6th International Competition on Plagiarism Detection. In CLEF 2014 Evaluation Labs and Workshop: Working Notes Papers.
Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of EMNLP-IJCNLP 2019 (pp. 3982–3992). Association for Computational Linguistics.
Rudwan, M. S. M., & Fonou-Dombeu, J. V. (2023). Hybridizing fuzzy string matching and machine learning for improved ontology alignment. Future Internet, 15(7), 229.
Salton, G., Wong, A., & Yang, C. S. (1975). A vector space model for automatic indexing. Communications of the ACM, 18(11), 613–620.
Santosa, F. (2022). Accuracy in identifying similarity levels in scientific articles using the Jaro Winkler algorithm. Jurnal Informasi dan Teknologi, 4(3), 142-147. https://doi.org/10.37034/jidt.v4i3.217
Singh, P. N., & Gowdar, T. P. (2021). Searching string in big-data: A better approach by applied machine learning. SN Computer Science, 2(3), 192.
Sudarma, M. (2025). Integrating LSTM and hybrid methods for automatic Balinese script transcription. Journal of Information Systems Engineering and Management, 10(9s), 709–720.
Winkler, W. E. (1990). String comparator metrics and enhanced decision rules in the Fellegi-Sunter model of record linkage. Proceedings of the Section on Survey Research Methods, American Statistical Association, 354–359.
DOI: https://doi.org/10.20527/cetj.v6i1.18697
Refbacks
- There are currently no refbacks.








