Acquiring Semantic Sibling Associations from Web Documents

Acquiring Semantic Sibling Associations from Web Documents

Marko Brunzel (German Research Center for Artificial Intelligence (DFKI GmbH), Germany) and Myra Spiliopoulou (Otto-von-Guericke-Universitat Magdeburg, Germany)
Copyright: © 2007 |Pages: 16
DOI: 10.4018/jdwm.2007100105
OnDemand PDF Download:


The automated discovery of relationships among terms contributes to the automation of the ontology engineering process and allows for sophisticated query expansion in information retrieval. While there are many findings on the identification of direct hierarchical relations among concepts, less attention has been paid on the discovery sibling terms. These are terms that share a common, a priori unknown parent such as co-hyponyms and co-meronyms. In this study, we present our results on the discovery of pairs or groups of sibling terms with XTREEM-SA (Xhtml TREE mining for sibling associations), an algorithm that extracts semantics from Web documents. While conventional methods process an appropriately prepared corpus, XTREEM-SA takes as input an arbitrary collection of Web documents on a given topic and finds sibling relations between terms in this corpus. It is thus independent of domain and language, does not require linguistic preprocessing, and does not rely on syntactic or other rules on text formation. We describe XTREEM-SA and evaluate it toward two reference ontologies. In this context, we also elaborate on the challenges of evaluating semantics extracted from the Web against handcrafted ontologies of high quality but possibly low coverage.

Complete Article List

Search this Journal:
Open Access Articles: Forthcoming
Volume 13: 4 Issues (2017): Forthcoming, Available for Pre-Order
Volume 12: 4 Issues (2016): 3 Released, 1 Forthcoming
Volume 11: 4 Issues (2015)
Volume 10: 4 Issues (2014)
Volume 9: 4 Issues (2013)
Volume 8: 4 Issues (2012)
Volume 7: 4 Issues (2011)
Volume 6: 4 Issues (2010)
Volume 5: 4 Issues (2009)
Volume 4: 4 Issues (2008)
Volume 3: 4 Issues (2007)
Volume 2: 4 Issues (2006)
Volume 1: 4 Issues (2005)
View Complete Journal Contents Listing