Language Models for XML Element Retrieval

Li, Rongmei; van der Weide, Theo

doi:10.1007/978-3-642-14556-8_11

Rongmei Li¹⁹ &
Theo van der Weide²⁰

Part of the book series: Lecture Notes in Computer Science ((LNISA,volume 6203))

Included in the following conference series:

International Workshop of the Initiative for the Evaluation of XML Retrieval

550 Accesses
2 Citations

Abstract

In this paper we describe our participation in the INEX 2009 ad-hoc track. We participated in all four retrieval tasks (thorough, focused, relevant-in-context, best-in-context) and report initial findings based on a single set of measure for all tasks. In this first participation, we test two ideas: (1) evaluate the performance of standard IR engines used in full document retrieval and XML element retrieval; (2) investigate if document structure can lead to more accurate and focused retrieval result. We find: 1) the full document retrieval outperforms the XML element retrieval using language model based on Dirichlet priors; 2) the element relevance score itself can be used to remove overlapping element results effectively.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 39.99; Price excludes VAT (USA)

Softcover Book: USD 54.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

Preview

Unable to display preview. Download preview PDF.

References

Blei, D.M., Ng, A.Y., Jordan, M.I.: Latent dirichlet allocation. J. Mach. Learn. Res. 3, 993–1022 (2003)
MATH Google Scholar
Zhai, C.X., Lafferty, J.: A Study of Smoothing Methods for Language Models Applied to Information Retrieval. ACM Trans. on Information Systems 22(2), 179–214 (2004)
Article Google Scholar
Schenkel, R., Suchanek, F.M., Kasneci, G.: YAWN: A Semantically Annotated Wikipedia XML Corpus. In: 12. GI-Fachtagung fr Datenbanksysteme in Business, Technologie und Web, Aachen, Germany (March 2007)
Google Scholar
Strohman, T., Metzler, D., Turtle, H., Croft, W.B.: Indri: A Language-model Based Search Engine for Complex Queries. In: Proceedings of ICIA (2005)
Google Scholar

Download references

Author information

Authors and Affiliations

University of Twente, Enschede, The Netherlands
Rongmei Li
Radboud University, Nijmegen, The Netherlands
Theo van der Weide

Authors

Rongmei Li
View author publications
You can also search for this author in PubMed Google Scholar
Theo van der Weide
View author publications
You can also search for this author in PubMed Google Scholar

Editor information

Editors and Affiliations

Faculty of Science and Technology, Queensland University of Technology, GPO Box 2434, 4001, Brisbane, Qld, Australia
Shlomo Geva
Archives and Information Studies/Humanities, University of Amsterdam, Turfdraagsterpad 9, 1012 XT, Amsterdam, The Netherlands
Jaap Kamps
Department of Computer Science, University of Otago, P.O. Box 56,, 9054, Dunedin, New Zealand
Andrew Trotman

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Li, R., van der Weide, T. (2010). Language Models for XML Element Retrieval. In: Geva, S., Kamps, J., Trotman, A. (eds) Focused Retrieval and Evaluation. INEX 2009. Lecture Notes in Computer Science, vol 6203. Springer, Berlin, Heidelberg. https://doi.org/10.1007/978-3-642-14556-8_11

Download citation

DOI: https://doi.org/10.1007/978-3-642-14556-8_11
Publisher Name: Springer, Berlin, Heidelberg
Print ISBN: 978-3-642-14555-1
Online ISBN: 978-3-642-14556-8
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics