Matching an XML Document against a Set of DTDs
Sources of XML documents are proliferating on the Web and documents are more and more frequently exchanged among sources. At the same time, there is an increasing need of exploiting database tools to manage this kind of data. An important novelty of XML is that information on document structures is available on the Web together with the document contents. However, in such an heterogeneous environment as the Web, it is not reasonable to assume that XML documents that enter a source always conform to a predefined DTD in the source. In this paper we address the problem of document classification by proposing a metric for quantifying the structural similarity between an XML document and a DTD. Based on such notion, we propose an approach to match a document entering a source against the set of DTDs available in the source, determining whether a DTD exists similar enough to the document.
Unable to display preview. Download preview PDF.
- 1.E. Bertino, G. Guerrini, I. Merlo, and M. Mesiti. An Approach to Classify Semi-Structured Objects. In Proc. European Conf. on Object-Oriented Programming, LNCS 1628, pp. 416–440, 1999.Google Scholar
- 2.E. Bertino, G. Guerrini, and M. Mesiti. Measuring the Structural Similarity among XML Documents and DTDs, 2001. http://www.disi.unige.it/person/MesitiM.
- 3.S. Castano, V. De Antonellis, M. G. Fugini, and B. Pernici. Conceptual Schema Analysis: Techniques and Applications. ACM Transactions on Database Systems, 23(3):286–333, Sept. 1998.Google Scholar
- 4.M. N. Garofalakis, A. Gionis, R. Rastogi, S. Seshadri, and K. Shim. XTRACT: A System for Extracting Document Type Descriptors from XML Documents. In Proc. of Int’l Conf. on Management of Data, pp. 165–176, 2000.Google Scholar
- 5.A. Miller. WordNet: A Lexical Database for English. Communications of the ACM, 38(11):39–41, Nov. 1995.Google Scholar
- 6.T. Milo and S. Zohar. Using Schema Matching to Simplify Heterogeneous Data Translation. In Proc. of Int’l Conf. on Very Large Data Bases, pp. 122–133, 1998.Google Scholar
- 7.S. Nestorov, S. Abiteboul, and R. Motwani. Extracting Schema from Semistructured Data. In Proc. of Int’l Conf. on Management of Data, pp. 295–306, 1998.Google Scholar
- 9.W3C. Document Object Model (DOM), 1998.Google Scholar
- 10.W3C. Extensible Markup Language 1.0, 1998.Google Scholar