A Two-Phase Sampling Technique to Improve the Accuracy of Text Similarities in the Categorisation of Hidden Web Databases

Hedley, Yih-Ling; Younas, Muhammad; James, Anne; Sanderson, Mark

doi:10.1007/978-3-540-30480-7_54

Yih-Ling Hedley²¹,
Muhammad Younas²¹,
Anne James²¹ &
…
Mark Sanderson²²

Part of the book series: Lecture Notes in Computer Science ((LNCS,volume 3306))

Included in the following conference series:

International Conference on Web Information Systems Engineering

1168 Accesses

Abstract

The larger amount of high quality and specialised information on the Web is stored in document databases, which is not indexed by general-purpose search engines such as Google and Yahoo. Such information is dynamically generated as a result of submitting queries to databases – which are referred to as Hidden Web databases. This paper presents a Two-Phase Sampling (2PS) technique that detects Web page templates from the randomly sampled documents of a database. It generates terms and frequencies that summarise the database content with improved accuracy. We then utilise such statistics to improve the accuracy of text similarity computation in categorisation. Experimental results show that 2PS effectively eliminates terms contained in Web page templates, and generates terms and frequencies with improved accuracy. We also demonstrate that 2PS improves the accuracy of text similarity computation required in the process of database categorisation.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 84.99; Price excludes VAT (USA)

Softcover Book: USD 109.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

Preview

Unable to display preview. Download preview PDF.

References

Bergman, M.K.: The Deep Web: Surfacing Hidden Value. Appeared in The Journal of Electronic Publishing from the University of Michigan (2001), http://www.press.umich.edu/jep/07-01/bergman.html (retrieved August 10, 2004)
Callan, J., Connell, M.: Query-Based Sampling of Text Databases. ACM Transactions on Information Systems 19(2), 97–130 (2001)
Article Google Scholar
Fravano, L., Change, K., Garcia-Molina, H., Paepcke, A.: STARTS Stanford Proposal for Internet Meta-Searching. In: Proceedings of the ACM-SIGMOD International Conference on Management of Data (1997)
Google Scholar
Gravano, L., Ipeirotis, P.G., Sahami, M.: QProber: A System for Automatic Classification of Hidden-Web Databases. ACM Transactions on Information Systems (TOIS) 21(1) (2003)
Google Scholar
Heß, M., Drobnik, O.: Clustering Specialised Web-databases by Exploiting Hyperlinks. In: Proceedings of the Second Asian Digital Library Conference (1999)
Google Scholar
Hedley, Y.L., Younas, M., James, A., Sanderson, M.: Query-Related Data Extraction of Hidden Web Documents. In: Proceedings of SIGIR (2004)
Google Scholar
Lin, K.I., Chen, H.: Automatic Information Discovery from the Invisible Web. In: International Conference on Information Technology: Coding and Computing (2002)
Google Scholar
Meng, W., Wang, W., Sun, H., Yu, C.: Concept Hierarchy Based Text Database Categorization. International Journal on Knowledge and Information Systems 4(2), 132–150 (2002)
Article Google Scholar
Salton, G., McGill, M.: Introduction to Modern Information Retrieval. McCraw-Hill, New York (1983)
Google Scholar
Sugiura, A., Etzioni, O.: Query Routing for Web Search Engines: Architecture and Experiment. In: 9th WWW Conference (2000)
Google Scholar

Download references

Author information

Authors and Affiliations

School of Mathematical and Information Sciences, Coventry University, Priory Street, Coventry, CV1 5FB, UK
Yih-Ling Hedley, Muhammad Younas & Anne James
Department of Information Studies, University of Sheffield, Regent Court, 211 Portobello St, Sheffield, S1 4DP, UK
Mark Sanderson

Authors

Yih-Ling Hedley
View author publications
You can also search for this author in PubMed Google Scholar
Muhammad Younas
View author publications
You can also search for this author in PubMed Google Scholar
Anne James
View author publications
You can also search for this author in PubMed Google Scholar
Mark Sanderson
View author publications
You can also search for this author in PubMed Google Scholar

Editor information

Editors and Affiliations

School of ITEE, The University of Queensland, Australia
Xiaofang Zhou
Database Systems Research and Development Center, University of Florida, P.O. Box 116125, 470 CSE, 32601-6125, Gainesville, FL, USA
Stanley Su
INFOLAB, Dept. of Information Systems and Management, Tilburg University, The Netherlands
Mike P. Papazoglou
Polish-Japanese Institute of Information Technology, Faculty of IT, Ul. Koszykowa 86, 02-008, Warsaw, Poland
Maria Elzbieta Orlowska
Rutherford Appleton Laboratory, Science and Technology Facilities Council, Harwell Science and Innovation Campus, OX11 0QX, Didcot, UK
Keith Jeffery

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Hedley, YL., Younas, M., James, A., Sanderson, M. (2004). A Two-Phase Sampling Technique to Improve the Accuracy of Text Similarities in the Categorisation of Hidden Web Databases. In: Zhou, X., Su, S., Papazoglou, M.P., Orlowska, M.E., Jeffery, K. (eds) Web Information Systems – WISE 2004. WISE 2004. Lecture Notes in Computer Science, vol 3306. Springer, Berlin, Heidelberg. https://doi.org/10.1007/978-3-540-30480-7_54

Download citation

DOI: https://doi.org/10.1007/978-3-540-30480-7_54
Publisher Name: Springer, Berlin, Heidelberg
Print ISBN: 978-3-540-23894-2
Online ISBN: 978-3-540-30480-7
eBook Packages: Springer Book Archive

Publish with us

Policies and ethics