Abstract
Sources of XML documents are proliferating on the Web and documents are more and more frequently exchanged among sources. At the same time, there is an increasing need of exploiting database tools to manage this kind of data. An important novelty of XML is that information on document structures is available on the Web together with the document contents. However, in such an heterogeneous environment as the Web, it is not reasonable to assume that XML documents that enter a source always conform to a predefined DTD in the source. In this paper we address the problem of document classification by proposing a metric for quantifying the structural similarity between an XML document and a DTD. Based on such notion, we propose an approach to match a document entering a source against the set of DTDs available in the source, determining whether a DTD exists similar enough to the document.
Access this chapter
Tax calculation will be finalised at checkout
Purchases are for personal use only
Preview
Unable to display preview. Download preview PDF.
References
E. Bertino, G. Guerrini, I. Merlo, and M. Mesiti. An Approach to Classify Semi-Structured Objects. In Proc. European Conf. on Object-Oriented Programming, LNCS 1628, pp. 416–440, 1999.
E. Bertino, G. Guerrini, and M. Mesiti. Measuring the Structural Similarity among XML Documents and DTDs, 2001. http://www.disi.unige.it/person/MesitiM.
S. Castano, V. De Antonellis, M. G. Fugini, and B. Pernici. Conceptual Schema Analysis: Techniques and Applications. ACM Transactions on Database Systems, 23(3):286–333, Sept. 1998.
M. N. Garofalakis, A. Gionis, R. Rastogi, S. Seshadri, and K. Shim. XTRACT: A System for Extracting Document Type Descriptors from XML Documents. In Proc. of Int’l Conf. on Management of Data, pp. 165–176, 2000.
A. Miller. WordNet: A Lexical Database for English. Communications of the ACM, 38(11):39–41, Nov. 1995.
T. Milo and S. Zohar. Using Schema Matching to Simplify Heterogeneous Data Translation. In Proc. of Int’l Conf. on Very Large Data Bases, pp. 122–133, 1998.
S. Nestorov, S. Abiteboul, and R. Motwani. Extracting Schema from Semistructured Data. In Proc. of Int’l Conf. on Management of Data, pp. 295–306, 1998.
A. Tversky. Features of Similarity. J. of Psychological Review, 84(4):327–352, 1977.
W3C. Document Object Model (DOM), 1998.
W3C. Extensible Markup Language 1.0, 1998.
Author information
Authors and Affiliations
Editor information
Editors and Affiliations
Rights and permissions
Copyright information
© 2002 Springer-Verlag Berlin Heidelberg
About this paper
Cite this paper
Bertino, E., Guerrini, G., Mesiti, M. (2002). Matching an XML Document against a Set of DTDs. In: Hacid, MS., Raś, Z.W., Zighed, D.A., Kodratoff, Y. (eds) Foundations of Intelligent Systems. ISMIS 2002. Lecture Notes in Computer Science(), vol 2366. Springer, Berlin, Heidelberg. https://doi.org/10.1007/3-540-48050-1_45
Download citation
DOI: https://doi.org/10.1007/3-540-48050-1_45
Published:
Publisher Name: Springer, Berlin, Heidelberg
Print ISBN: 978-3-540-43785-7
Online ISBN: 978-3-540-48050-1
eBook Packages: Springer Book Archive