Integrating Feature Analysis and Background Knowledge to Recommend Similarity Functions

Ryu, Seung Hwan; Benatallah, Boualem

doi:10.1007/978-3-642-35063-4_52

Seung Hwan Ryu²⁰ &
Boualem Benatallah²⁰

Part of the book series: Lecture Notes in Computer Science ((LNISA,volume 7651))

Included in the following conference series:

International Conference on Web Information Systems Engineering

2537 Accesses

Abstract

Existing approaches in similarity analysis is little concerned with the right choice of similarity functions. We present an approach for suggesting which similarity functions (e.g., edit distance) are most appropriate for a given similarity search task. We identify data features (e.g., misspellings) that are considerable when choosing similarity functions. We also introduce the concept of similarity function background knowledge that associates data features with similarity functions, and apply the knowledge to recommend suitable similarity functions.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 39.99; Price excludes VAT (USA)

Softcover Book: USD 54.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

Preview

Unable to display preview. Download preview PDF.

References

Tailor: A record linkage tool box. In: Proceedings of the 18th International Conference on Data Engineering, ICDE 2002. IEEE Computer Society (2002)
Google Scholar
Bilenko, M., Mooney, R.J.: Adaptive duplicate detection using learnable string similarity measures. In: KDD, pp. 39–48. ACM (2003)
Google Scholar
Bilenko, M., Mooney, R.J., Cohen, W.W., Ravikumar, P.D., Fienberg, S.E.: Adaptive name matching in information integration. IEEE Int. Sys. (2003)
Google Scholar
Chaudhuri, S., Chen, B.-C., Ganti, V., Kaushik, R.: Example-driven design of efficient record matching queries. In: VLDB, pp. 327–338 (2007)
Google Scholar
Christen, P.: A comparison of personal name matching: Techniques and practical issues. In: ICDM Workshops, pp. 290–294 (2006)
Google Scholar
Dong, X., Halevy, A.Y., Madhavan, J.: Reference reconciliation in complex information spaces. In: SIGMOD Conference, pp. 85–96 (2005)
Google Scholar
Elmagarmid, A.K., Ipeirotis, P.G., Verykios, V.S.: Duplicate record detection: A survey. IEEE Trans. Knowl. Data Eng. 19(1), 1–16 (2007)
Article Google Scholar
Hall, P.A.V., Dowling, G.R.: Approximate string matching. ACM Comput. Surv. 12, 381–402 (1980)
Article MathSciNet Google Scholar
Köpcke, H., Thor, A., Rahm, E.: Learning-based approaches for matching web data entities. IEEE Internet Computing 14(4), 23–31 (2010)
Article Google Scholar
Lange, D., Naumann, F.: Efficient similarity search: arbitrary similarity measures, arbitrary composition. In: CIKM, pp. 1679–1688 (2011)
Google Scholar
Ryu, S.H., Benatallah, B., Paik, H.-Y., Kim, Y.S., Compton, P.: Similarity Function Recommender Service Using Incremental User Knowledge Acquisition. In: Kappel, G., Maamar, Z., Motahari-Nezhad, H.R. (eds.) ICSOC 2011. LNCS, vol. 7084, pp. 219–234. Springer, Heidelberg (2011)
Chapter Google Scholar
Sahar, S.: Interestingness via what is not interesting. In: KDD, pp. 332–336 (1999)
Google Scholar
Tejada, S., Knoblock, C.A., Minton, S.: Learning domain-independent string transformation weights for high accuracy object identification. In: KDD (2002)
Google Scholar
Wang, J., Li, G., Feng, J.: Fast-join: An efficient method for fuzzy token matching based string similarity join. In: ICDE, pp. 458–469 (2011)
Google Scholar

Download references

Author information

Authors and Affiliations

School of Computer Science & Engineering, University of New South Wales, Sydney, NSW, 2051, Australia
Seung Hwan Ryu & Boualem Benatallah

Authors

Seung Hwan Ryu
View author publications
You can also search for this author in PubMed Google Scholar
Boualem Benatallah
View author publications
You can also search for this author in PubMed Google Scholar

Editor information

Editors and Affiliations

School of Computer Science, Fudan University, 825 Zhangheng Rd., Shanghai, 201203, China
X. Sean Wang
Department of Computer Science, College of Engineering, Science and Engineering Offices, The University of Illinois at Chicago, 851 South Morgan Street (M/C 152), 60607-7053, Chicago, Illinois, USA
Isabel Cruz
Department of Informatics and Telecommunications, University of Athens, GR15784, Ilisia, Athens, Greece
Alex Delis
Centre for Applied Informatics, Victoria University, PO Box 14428, 8001, Melbourne, VIC, Australia
Guangyan Huang

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Ryu, S.H., Benatallah, B. (2012). Integrating Feature Analysis and Background Knowledge to Recommend Similarity Functions. In: Wang, X.S., Cruz, I., Delis, A., Huang, G. (eds) Web Information Systems Engineering - WISE 2012. WISE 2012. Lecture Notes in Computer Science, vol 7651. Springer, Berlin, Heidelberg. https://doi.org/10.1007/978-3-642-35063-4_52

Download citation

DOI: https://doi.org/10.1007/978-3-642-35063-4_52
Publisher Name: Springer, Berlin, Heidelberg
Print ISBN: 978-3-642-35062-7
Online ISBN: 978-3-642-35063-4
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics