Accessing the Deep Web with Keywords: A Foundational Approach
- 555 Downloads
The Deep Web is constituted by data that are generated dynamically as the result of interactions with Web pages. The problem of accessing Deep Web data presents many challenges: it has been shown that answering even simple queries on such data requires the execution of recursive query plans. There is a gap between the theoretical understanding of this problem and the practical approaches to it. The main reason behind this is that the problem is to be studied by considering the database as part of the input, but queries can be processed by accessing data according to limitations, expressed as so-called access patterns. In this paper we embark on the task of closing the above gap by giving a precise definition that reflects the practical nature of accessing Deep Web data sources. In particular, we define the problem of querying Deep Web sources with keywords. We describe two scenarios: in the first, called unrestricted, there query answering algorithm has full access to the data; in the second, called restricted, the algorithm can access the data only according to the access patterns. We formalise the associated decision problem associated to that of query answering in the Deep Web, explaining its relevance in both the aforementioned scenarios. We then present some complexity results.
KeywordsUnrestricted Case Initial Keyword Atomic Queries Conjunctive Queries (CQ) Abstract Domain
This work was supported by the EU COST Action IC1302 KEYSTONE. Andrea Calì acknowledges partial support by the EPSRC project “Logic-based Integration and Querying of Unindexed Data” (EP/E010865/1).
- 3.Calì, A., Martinenghi, D.: Querying data under access limitations. In: Proceedings of ICDE (2008)Google Scholar
- 4.Calì, A., Martinenghi, D., Razgon, I., Ugarte, M.: Querying the deep web: back to the foundations. In: Proceedings of AMW (2017). To appearGoogle Scholar
- 5.Calì, A., Razgon, I.: Complexity of conjunctive query answering under access limitations (preliminary report). In: Proceedings of SEBD (2014)Google Scholar
- 6.Chang, K.C.-C., He, B., Zhang, Z.: Toward large scale integration: building a metaquerier over databases on the web. In: Proceedings of CIDR (2005)Google Scholar
- 8.Li, C., Chang, E.: Query planning with limited source capabilities. In: Proceedings of ICDE (2000)Google Scholar
- 9.Madhavan, J., Afanasiev, L., Antova, L., Halevy, A.Y.: Harnessing the deep web: present and future. In: Proceedings of CIDR (2009)Google Scholar