Extended Language Models for XML Element Retrieval

Li, Rongmei; van der Weide, Theo

doi:10.1007/978-3-642-23577-1_8

Rongmei Li²⁰ &
Theo van der Weide²⁰

Part of the book series: Lecture Notes in Computer Science ((LNISA,volume 6932))

Included in the following conference series:

International Workshop of the Initiative for the Evaluation of XML Retrieval

413 Accesses
1 Citations

Abstract

In this paper we describe our participation in the INEX 2010 ad-hoc track. We participated in three retrieval tasks (restricted focused task, relevant-in-context, restricted relevant-in-context) and report our findings based on a single set of measure for all tasks. In this year’s participation, we evaluate the performance of the standard language model that is more focused on a fixed number of relevant characters than on relevant paragraphs. Our findings are: 1) the simplest language model for document retrieval performs relatively well in the restricted focused task when using a fixed offset that is close to the average character distance from the beginning of a document to its main content; 2) a good result of document ranking does improve the performance of snippet retrieval; 3) stemming and stopword removal can further boost performance.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Subscribe and save

Springer+ Basic

$34.99 /Month

Get 10 units per month
Download Article/Chapter or eBook
1 Unit = 1 Article or 1 Chapter
Cancel anytime

Buy Now

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 39.99; Price excludes VAT (USA)

Softcover Book: USD 54.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

Preview

Unable to display preview. Download preview PDF.

Information Retrieval in XML Document: State of the Art

Adapting LLMs for Efficient, Personalized Information Retrieval: Methods and Implications

From XML Retrieval to Semantic Search and Beyond

References

Li, R.M., van der Weide, T.P.: Language Models for XML Element Retrieval. In: Proceedings of INEX (2009)
Google Scholar
Zhai, C.X., Lafferty, J.: A Study of Smoothing Methods for Language Models Applied to Information Retrieval. ACM Trans. on Information Systems. 22(2), 179–214 (2004)
Article Google Scholar
Schenkel, R., Suchanek, F.M., Kasneci, G.: YAWN: A Semantically Annotated Wikipedia XML Corpus. In: 12. GI-Fachtagung für Datenbanksysteme in Business, Technologie und Web, Aachen, Germany (March 2007)
Google Scholar
Strohman, T., Metzler, D., Turtle, H., Croft, W.B.: Indri: A Language-model Based Search Engine for Complex Queries. In: Proceedings of ICIA (2005)
Google Scholar

Download references

Author information

Authors and Affiliations

Radboud University, Nijmegen, The Netherlands
Rongmei Li & Theo van der Weide

Authors

Rongmei Li
View author publications
You can also search for this author in PubMed Google Scholar
Theo van der Weide
View author publications
You can also search for this author in PubMed Google Scholar

Editor information

Editors and Affiliations

Faculty of Science and Technology, Queensland University of Technology, GPO Box 2434, Qld 4001, Brisbane, Australia
Shlomo Geva
Archives and Information Studies/Humanities, University of Amsterdam, Turfdraagsterpad 9, 1012XT, Amsterdam, The Netherlands
Jaap Kamps
Multimodal Computing and Interaction, Saarland University, 66123, Saarbrücken, Germany
Ralf Schenkel
Department of Computer Science, University of Otago, P.O. Box 56, 9054, Dunedin, New Zealand
Andrew Trotman

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Li, R., van der Weide, T. (2011). Extended Language Models for XML Element Retrieval. In: Geva, S., Kamps, J., Schenkel, R., Trotman, A. (eds) Comparative Evaluation of Focused Retrieval. INEX 2010. Lecture Notes in Computer Science, vol 6932. Springer, Berlin, Heidelberg. https://doi.org/10.1007/978-3-642-23577-1_8

Download citation

DOI: https://doi.org/10.1007/978-3-642-23577-1_8
Publisher Name: Springer, Berlin, Heidelberg
Print ISBN: 978-3-642-23576-4
Online ISBN: 978-3-642-23577-1
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics