Skip to main content
Log in

Retrieval effectiveness of an ontology-based model for information selection

  • Published:
The VLDB Journal Aims and scope Submit manuscript

Abstract.

Technology in the field of digital media generates huge amounts of nontextual information, audio, video, and images, along with more familiar textual information. The potential for exchange and retrieval of information is vast and daunting. The key problem in achieving efficient and user-friendly retrieval is the development of a search mechanism to guarantee delivery of minimal irrelevant information (high precision) while insuring relevant information is not overlooked (high recall). The traditional solution employs keyword-based search. The only documents retrieved are those containing user-specified keywords. But many documents convey desired semantic information without containing these keywords. This limitation is frequently addressed through query expansion mechanisms based on the statistical co-occurrence of terms. Recall is increased, but at the expense of deteriorating precision. One can overcome this problem by indexing documents according to context and meaning rather than keywords, although this requires a method of converting words to meanings and the creation of a meaning-based index structure. We have solved the problem of an index structure through the design and implementation of a concept-based model using domain-dependent ontologies. An ontology is a collection of concepts and their interrelationships that provide an abstract view of an application domain. With regard to converting words to meaning, the key issue is to identify appropriate concepts that both describe and identify documents as well as language employed in user requests. This paper describes an automatic mechanism for selecting these concepts. An important novelty is a scalable disambiguation algorithm that prunes irrelevant concepts and allows relevant ones to associate with documents and participate in query generation. We also propose an automatic query expansion mechanism that deals with user requests expressed in natural language. This mechanism generates database queries with appropriate and relevant expansion through knowledge encoded in ontology form. Focusing on audio data, we have constructed a demonstration prototype. We have experimentally and analytically shown that our model, compared to keyword search, achieves a significantly higher degree of precision and recall. The techniques employed can be applied to the problem of information selection in all media types.

This is a preview of subscription content, log in via an institution to check access.

Access this article

Price excludes VAT (USA)
Tax calculation will be finalised during checkout.

Instant access to the full article PDF.

Similar content being viewed by others

References

  1. Adali S, Candan KS, Chen S, Erol K, Subrahmanian VS (1996) Advanced video information system: data structures and query processing. ACM Multimedia Sys J 4:172-186

    Article  Google Scholar 

  2. Agirre E, Ansa O, Hovy E, Martinez D (2000) Enriching very large ontologies using the WWW. In: Proceedings of the ontology learning workshop, ECAI, Berlin, Germany, August 2000

  3. Arons B (1993) SpeechSkimmer: interactively skimming recorded speech. In: Proceedings of the ACM symposium on user interface software and technology, Atlanta, November 1993, pp 187-196

  4. Aslan G, McLeod D (1999) Semantic heterogeneity resolution in federated database by metadata implantation and stepwise evolution. VLDB J 18(2):120-132

    Article  Google Scholar 

  5. Baeza R, Neto B (1999) Modern information retrieval. ACM Press, New York, Addison-Wesley, Reading, MA

  6. Bunge M (1977) Treatise on basic philosophy. Ontology I. The furniture of the world, vol 3. Reidel, Boston

  7. Gibbs S, Breitender C, Tsichritzis D (1994) Data modeling of time based media. In: Proceedings of ACM SIGMOD international conference on management of data, Minneapolis, June 1994, pp 91-102

  8. Gonzalo J, Verdejo F, Chugur I, Cigarran J (1998) Indexing with WordNet synsets can improve text retrieval. In: Proceedings of the Coling-ACL’98 workshop: usage of WordNet in natural language processing systems, Montreal, August 1998, pp 38-44

  9. Gruber TR (1993) A translation approach to portable ontology specifications. Knowl Acquisition 5(2):199-220

    Article  Google Scholar 

  10. Guarino N, Masolo C, Vetere G (1999) OntoSeek: content-based access to the Web. IEEE Intell Sys 14(3):70-80

    Article  Google Scholar 

  11. Hauptmann AG (1995) Speech recognition in the informedia digital video library: uses and limitations. In: Proceedings of the 7th IEEE international conference on tools with AI, Washington, DC, November 1995

  12. Hirschberg J, Grosz B (1991) Intonational features of local and global discourse. In: Proceedings of the speech and natural language workshop, Harriman, NY, February 1991, pp 23-26

  13. Hjelsvold R, Midstraum R (1994) Modeling and querying video data. In: Proceedings of the 20th international conference on very large databases (VLDB’94), Santiago, Chile, September 1994, pp 686-694

  14. Khan L, McLeod D (2000) Audio structuring and personalized retrieval using ontologies. In: Proceedings of IEEE advances in digital libraries, library of congress, Bethesda, MD, May 2000, pp 116-126

  15. Khan L, McLeod D (2000) Effective retrieval of audio information from annotated text using ontologies. In: Proceedings of the ACM SIGKDD workshop on multimedia data mining, Boston, August 2000, pp 37-45

  16. Labrou Y, Finin T (1999) Yahoo! as an ontology - using Yahoo! categories to describe documents. In: Proceedings of the 8th international conference on knowledge and information management (CIKM-99), Kansas City, MO, October 1999, pp 180-187

  17. Lenat DB (1995) Cyc: a large-scale investment in knowledge infrastructure. Commun ACM 38(11):33-38

    Article  Google Scholar 

  18. Miller G (1995) WordNet: a lexical database for English. Commun ACM 38(11):39-41

    Article  MATH  Google Scholar 

  19. Mitra P, Kersten M, Wiederhold G (2000) Graph-oriented model for articulation of ontology interdependencies. In: Proceedings of the 7th international conference on extending database technology (EDBT 2000), Konstanz, Germany, March 2000, pp 86-100

  20. Omoto E, Tanaka K (1993) OVID: design and implementation of a video-object database system. IEEE Trans Knowl Data Eng 5(4):629-643

    Article  Google Scholar 

  21. Peat HJ, Willett P (1991) The limitations of term co-occurrence data for query expansion in document retrieval systems. J ASIS 42(5):378-383

    Google Scholar 

  22. Rabiner LR, Schafer RW (1978) Digital processing of speech signals. Prentice-Hall, Upper Saddle River, NJ

  23. Salton G (1989) Automatic text processing. Addison-Wesley, Reading, MA

  24. Smeaton AF, Rijsbergen V (1993) The retrieval effects of query expansion on a feedback document retrieval system. Comput J 26(3):239-246

    Google Scholar 

  25. Swartout B, Patil R, Knight K, Ross T (1996) Toward distributed use of large-scale ontologies. In: Proceedings of the 10th workshop on knowledge acquisition for knowledge-based systems, Banff, Canada, 1996

  26. Using XML: Ontology and Conceptual Knowledge Markup Languages (1999) http://www.oasis-open.org/cover/xml.html

  27. Voorhees E (1994) Query expansion using lexical-semantic relations. In: Proceedings of the 17th annual international ACM SIGIR conference on research and development in information retrieval, Dublin, Ireland, July 1994, pp 61-69

  28. Wilcox LD, Bush MA (1992) Training and search algorithms for an interactive wordspotting system. In: Proceedings of the IEEE conference on acoustics, speech, and signal processing, San Francisco, vol 2, pp 97-100

  29. Woods W (1999) Conceptual indexing: a better way to organize knowledge. Technical report of Sun Microsystems

    Google Scholar 

Download references

Author information

Authors and Affiliations

Authors

Corresponding author

Correspondence to Latifur Khan.

Additional information

Received: 7 October 2002, Accepted: 20 May 2003, Published online: 30 September 2003

Edited by: E. Lochovsky

This research has been funded [or funded in part] by the Integrated Media Systems Center, a National Science Foundation Engineering Research Center, Cooperative Agreement No. EEC-9529152.

Rights and permissions

Reprints and permissions

About this article

Cite this article

Khan, L., McLeod, D. & Hovy, E. Retrieval effectiveness of an ontology-based model for information selection. VLDB 13, 71–85 (2004). https://doi.org/10.1007/s00778-003-0105-1

Download citation

  • Issue Date:

  • DOI: https://doi.org/10.1007/s00778-003-0105-1

Keywords:

Navigation