Semantic Document Processing Using Wikipedia as a Knowledge Base

Witten, Ian H.

doi:10.1007/978-3-642-14556-8_3

Ian H. Witten¹⁹

Part of the book series: Lecture Notes in Computer Science ((LNISA,volume 6203))

Included in the following conference series:

International Workshop of the Initiative for the Evaluation of XML Retrieval

571 Accesses

Abstract

Wikipedia is a goldmine of information; not just for its many readers, but also for the growing community of researchers who recognize it as a resource of exceptional scale and utility. It represents a vast investment of manual effort and judgment: a huge, constantly evolving tapestry of concepts and relations that is being applied to a host of tasks.

This talk will introduce the process of “wikification”; that is, automatically and judiciously augmenting a plain-text document with pertinent hyperlinks to Wikipedia articles – as though the document were itself a Wikipedia article. This amounts to a new semantic representation of text in terms of the salient concepts it mentions, where “concept” is equated to “Wikipedia article.” Wikification is a useful process in itself, adding value to plain text documents. More importantly, it supports new methods of document processing.

I first describe how Wikipedia can be used to determine semantic relatedness, and then introduce a new, high-performance method of wikification that exploits Wikipedia’s 60 M internal hyperlinks for relational information and their anchor texts as lexical information, using simple machine learning. I go on to discuss applications to knowledge-based information retrieval, topic indexing, document tagging, and document clustering. Some of these perform at human levels. For example, on CiteULike data, automatically extracted tags are competitive with tag sets assigned by the best human taggers, according to a measure of consistency with other human taggers.

Although this work is based on English it involves no syntactic parsing, and the techniques are largely language independent. The talk will include live demos.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Subscribe and save

Springer+ Basic

$34.99 /Month

Get 10 units per month
Download Article/Chapter or eBook
1 Unit = 1 Article or 1 Chapter
Cancel anytime

Buy Now

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 39.99; Price excludes VAT (USA)

Softcover Book: USD 54.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

Creating Textual Corpora Based on Wikipedia and Knowledge Graphs

YAGO: A Multilingual Knowledge Base from Wikipedia, Wordnet, and Geonames

Building Wikipedia N-grams with Apache Spark

Author information

Authors and Affiliations

Department of Computer Science, University of Waikato, New Zealand
Ian H. Witten

Authors

Ian H. Witten
View author publications
You can also search for this author in PubMed Google Scholar

Editor information

Editors and Affiliations

Faculty of Science and Technology, Queensland University of Technology, GPO Box 2434, 4001, Brisbane, Qld, Australia
Shlomo Geva
Archives and Information Studies/Humanities, University of Amsterdam, Turfdraagsterpad 9, 1012 XT, Amsterdam, The Netherlands
Jaap Kamps
Department of Computer Science, University of Otago, P.O. Box 56,, 9054, Dunedin, New Zealand
Andrew Trotman

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Witten, I.H. (2010). Semantic Document Processing Using Wikipedia as a Knowledge Base. In: Geva, S., Kamps, J., Trotman, A. (eds) Focused Retrieval and Evaluation. INEX 2009. Lecture Notes in Computer Science, vol 6203. Springer, Berlin, Heidelberg. https://doi.org/10.1007/978-3-642-14556-8_3

Download citation

DOI: https://doi.org/10.1007/978-3-642-14556-8_3
Publisher Name: Springer, Berlin, Heidelberg
Print ISBN: 978-3-642-14555-1
Online ISBN: 978-3-642-14556-8
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics

Semantic Document Processing Using Wikipedia as a Knowledge Base

Abstract

Access this chapter

Subscribe and save

Buy Now

Similar content being viewed by others

Creating Textual Corpora Based on Wikipedia and Knowledge Graphs

YAGO: A Multilingual Knowledge Base from Wikipedia, Wordnet, and Geonames

Building Wikipedia N-grams with Apache Spark

Author information

Authors and Affiliations

Editor information

Editors and Affiliations

Rights and permissions

Copyright information

About this paper

Cite this paper

Download citation

Publish with us

Subscribe and save

Buy Now

Navigation

Semantic Document Processing Using Wikipedia as a Knowledge Base

Abstract

Access this chapter

Subscribe and save

Buy Now

Similar content being viewed by others

Creating Textual Corpora Based on Wikipedia and Knowledge Graphs

YAGO: A Multilingual Knowledge Base from Wikipedia, Wordnet, and Geonames

Building Wikipedia N-grams with Apache Spark

Author information

Authors and Affiliations

Editor information

Editors and Affiliations

Rights and permissions

Copyright information

About this paper

Cite this paper

Download citation

Share this paper

Publish with us

Search

Navigation