Semantic Similarity of Product and Service Names in Portuguese
An Approach Based on Onto.PT
DOI:
https://doi.org/10.19153/cleiej.25.3.3Keywords:
Semantic Similarity, Portuguese, Ontology, Short Text, Onto.PTAbstract
The problem of conceptual comparison of names plays an important role in the field of natural language processing. In this task, the goal is to choose, among a set of names, which one refers to the same concept or object as a given input name. In this paper, we propose an algorithm for comparing names of products and services in Portuguese that takes account of the semantic information contained in the names. The semantic similarity between two names is calculated using information from Onto.PT, the largest public lexical ontology for the Portuguese language. Experiments were conducted on a dataset composed of 5,000 pairs of names of products and services in Portuguese. Our experimental results show that the algorithm based on Onto.PT is more effective than other well-known algorithms for name comparison, producing the highest recall and precision. Moreover, results also provide interesting insights into the advantages and disadvantages of using Onto.PT for assessing the semantic similarity of names and other kinds of short texts.
References
A. E. Monge and C. P. Elkan, “The field matching problem: Algorithms and applications,” in Proc. of the 2nd International Conference on Knowledge Discovery and Data Mining (KDD'96), Portland, US, Aug. 1996, pp. 267–270.
K. L. A. Branting, “A Comparative evaluation of name matching algorithms,” in Proc. of the 9th International Conference on Artificial Intelligence and Law (ICAIL’03), Edinburgh, Scotland, UK, Jun. 2003, pp. 224–232. [Online]. Available: https://doi.org/10.1145/1047788.1047837
F. M. Anuar, R. Setchi, and Y-K Lai, “Semantic Retrieval of Trademarks based on Conceptual Similarity,” IEEE Trans. on System, Man, and Cybernetics: Systems, vol 46, no. 2, pp. 220-233, Feb. 2016. [Online]. Available: https://doi.org/10.1109/TSMC.2015.2421878
N. Gali, R. Mariescu-Istodor, D. Hostettler, and P. Fränti, “Framework for Syntactic String Similarity Measures,” Expert Systems with Applications, vol 129, pp. 169-185, Sep. 2019. [Online]. Available: https://doi.org/10.1016/j.eswa.2019.03.048
I. V. Levenshtein, “Binary Codes Capable of Correcting Deletions, Insertions and Reversals,” Cybernetics and Control Theory, vol 10, no. 8, pp. 707-710, 1966.
P. Christen, “A comparison of personal name matching: Techniques and practical issues,” in Proc. of the IEEE 6th Data Mining Workshop (ICDMW’06), Hong Kong, China, Dec. 2006, pp. 290–294. [Online]. Available: https://doi.org/10.1109/ICDMW.2006.2
C. A. Davis Jr. and E. Salles, “Approximate string matching for geographic names and personal names,” in Proc. of the IX Brazilian Symposium on Geoinformatics, Campos do Jordão, Brazil, Nov. 2007, pp. 49–60.
G. Recchia and M. Louwerse, “A comparison of string similarity measures for toponym matching,” in Proc. of the 1st ACM SIGSPATIAL International Workshop on Computational Models of Place , Orlando, USA, Nov. 2013, pp. 54–61. [Online]. Available: https://doi.org/10.1145/2534848.2534850
H. G. Oliveira and P. Gomes, “ECO and Onto.PT: A Flexible Approach for Creating a Portuguese Wordnet Automatically,” Language Resources and Evaluation, vol 48, no. 2, pp. 373-393, Jun. 2014. [Online]. Available: https://doi.org/10.1007/s10579-013-9249-9
D. Lin, “An information-theoretic definition of similarity,” in Proc. of the 15th International Conference on Machine Learning (ICML’98), San Francisco, USA, Jul. 1998, pp. 296–304.
M. A. Jaro, “Advances in Record-Linkage Methodology as Applied to Matching the 1985 Census of Tampa, Florida,” Journal of the American Statistical Association, vol 84, no. 406, pp. 414-420, 1989. [Online]. Available: https://doi.org/10.1080/01621459.1989.10478785
J. Leskovec, A. Rajaraman, and J. Ullman, Mining of Massive Datasets. 3rd ed., New York: Cambridge University Press, 2020.
L. Gravano, et al., “Using q-grams in a DBMS for Approximate String Processing,” IEEE Data Engineering Bulletin, vol 24, no. 4, pp. 28-34, Dec. 2001.
W. E. Winkler, “Advanced methods for record linkage”, in Proc. of the Survey Research Methods Section, American Statistical Association, p. 467-472, 1994.
R. Sinoara, J. Antunes, and S. O. Rezende, “Text Mining and Semantics: A Systematic Mapping Study,” Journal of the Brazilian Computer Society, vol 23, no. 9, pp. 1-20, Jun. 2017. [Online]. Available: https://doi.org/10.1186/s13173-017-0058-7
D. S. M. Gomes, F. C. Cordeiro, B. S. Consoli, N. L. Santos, V. P. Moreira, R. Vieira, S. Moraes, A. G. Evsukoff, “Portuguese Word Embeddings for the Oil and Gas Industry: Development and Evaluation,” Computers in Industry, vol 1024, pp. 1-14, Jan. 2021. [Online]. Available: https://doi.org/10.1016/j.compind.2020.103347
T. P. Meirelles, E. C. Gonçalves, and D. T. Gomes, “Pareamento de Nomes de Produtos e Serviços Utilizando Medidas de Similaridade Textual nos Níveis Alfabético, Léxico e Semântico,” Cadernos do IME – Série Informática, vol 129, pp. 104-117, Dec. 2021. [Online]. Available: http://dx.doi.org/10.12957/cadinf
A. S. Romualdo, L. Real, and H. M. Caseli, “Measuring brazilian portuguese product titles similarity using embeddings,” in Proc. of the XIII Symposium in Information and Human Language Technology (STIL’21), Porto Alegre, Brazil, Nov. 2021, pp. 121–132. [Online]. Available: https://doi.org/10.5753/stil.2021.17791
Y. Li, D. McLean, Z. A. Bandar, J. D. O'Shea, and K. Crockett, “Sentence similarity based on semantic nets and corpus statistics,” IEEE Transactions on Knowledge and Data Engineering, vol 18, no. 8, pp. 1138-1150, Jun. 2006. [Online]. Available: https://doi.org/10.1109/TKDE.2006.130
D. Croft, S. Coupland, J. Shell, and S. C. Brown, “A fast and efficient semantic short text measure,” in Proc. of the 13rd UK Workshop on Computational Intelligence (UKCI’13), Guildford, UK, Sep. 2013, pp. 221–227. [Online]. Available: https://doi.org/10.1109/UKCI.2013.6651309
K. R. Varshney, "Data Science of the people, for the people, by the people: A viewpoint on an emerging dichotomy," in Proc. of the Bloomberg Data for Good Exchange 2015 (D4GX'2015), New York, USA, Sep. 2015, pp. 1–6.
C. Fellbaum, WordNet: an electronic lexical database. 3rd ed., New York: Cambridge University Press, 2020.
Z. Wu and M. Palmes, “Verbs semantics and lexical selection,” in Proc. of the 32nd Annual Meeting of the Association for Computational Linguistics, New Mexico, USA, Jun. 1994, pp. 133–138. [Online]. Available: https://doi.org/10.3115/981732.981751
N. Gali, R. Mariescu-Istodor, and P. Fränti, “Similarity measures for title matching,” in Proc. of the 23rd International Conference on Pattern Recognition (ICPR’16), Cancún, Mexico, Dec. 2016, pp. 1549–1554. [Online]. Available: https://doi.org/10.1109/ICPR.2016.7899857
V. de Paiva, L. Real, H. G. Oliveira, A. Rademaker, C. Freitas, and A. Simões, “An overview of portuguese wordnets”, in Proc. of the 8th Global WordNet Conference (GWC’16), Bucharest, Romania, Jan. 2016, pp. 74-81.
A. Branco, S. Grilo, M. Bolrinha, C. Saedi, R. Branco, J. Silva, A. Querido, R. de Carvalho, R. Gaudio, M. Avelãs, and C. Pinto, “The MWN.PT wordnet for portuguese: projection, validation, cross-lingual alignment and distribution,” in Proc. of the 12th Conference on Language Resources and Evaluation (LREC’20), Marselle, France, Dec. 2020, pp. 4859-4866.
V. de Paiva, A. Rademaker, and G. Melo “OpenWordnet-pt: An open brazilian wordnet for reasoning”, in Proc. of the Cooling 2012: Demonstration Papers, Mumbai, India, Dec. 2012, pp. 353-360.
H. G. Oliveira, “CONTO.PT: Groundwork for the automatic creation of a fuzzy portuguese wordnet,” in Proc. of the 12th International Conference on the Computational Processing of Portuguese (PROPOR’16), Tomar, Portugal, Jul. 2016, pp. 283-295. [Online]. Available: https://doi.org/10.1007/978-3-319-41552-9_29
Onto.PT. Ontologia Lexical para o Português. 2022. http://ontopt.dei.uc.pt/
IBGE. POF – Consumer Expenditure Survey. 2022. https://www.ibge.gov.br/en/statistics/social/health/25610-pof-2017-2018-pof-en.html?=&t=o-que-e
IBGE. IPCA - Extended National Consumer Price Index. 2022. https://www.ibge.gov.br/en/statistics/economic/prices-and-costs/17129-extended-national-consumer-price-index.html?=&t=o-que-e
C. D. Manning, P. Raghavan, and H. Schutze, Introduction to Information Retrieval. New York: Cambridge University Press, 2008.
J. Baeza-Yates and B Ribeiro-Neto, Modern Information Retrieval: The Concepts and Technology Behind Search. 2nd edition, New York: Addison-Wesley Professional, 2011.
strsimpy 0.2.1. python-string-similarity. 2022. https://pypi.org/project/strsimpy/
D. Jurafsky and J. H. Martin, Speech and Language Processing, 3rd edition (draft), 2021. [Online]. Available: https://web.stanford.edu/~jurafsky/slp3/
N. S. Hartmann, “Solo Queue at ASSIN: Combinando Abordagens Tradicionais e Emergentes,” Linguamática, vol 8, no. 2, pp. 59-64, Dec. 2016.
J.Yang, Y. Li, G. Gao, and Y. Zhang. “Measuring the Short Text Similarity Based on Semantic and Syntactic Information,” Future Generation Computer Systems, vol 114, pp. 169–180, Jan 2021. [Online]. Available: https://doi.org/10.1016/j.future.2020.07.043
Downloads
Published
Issue
Section
License
Copyright (c) 2023 Eduardo Gonçalves

This work is licensed under a Creative Commons Attribution 4.0 International License.
CLEIej is supported by its home institution, CLEI, and by the contribution of the Latin American and international researchers community, and it does not apply any author charges whatsoever for submitting and publishing. Since its creation in 1998, all contents are made publicly accesibly. The current license being applied is a (CC)-BY license (effective October 2015; between 2011 and 2015 a (CC)-BY-NC license was used).