Text Clustering with Limited User Feedback Under Local Metric Learning

Huang, Ruizhang; Zhang, Zhigang; Lam, Wai

doi:10.1007/11880592_11

Ruizhang Huang²⁰,
Zhigang Zhang²⁰ &
Wai Lam²⁰

Part of the book series: Lecture Notes in Computer Science ((LNISA,volume 4182))

Included in the following conference series:

Asia Information Retrieval Symposium

976 Accesses
1 Citations

Abstract

This paper investigates the idea of incorporating incremental user feedbacks and a small amount of sample documents for some, not necessarily all, clusters into text clustering. For the modeling of each cluster, we make use of a local weight metric to reflect the importance of the features for a particular cluster. The local weight metric is learned using both the unlabeled data and the constraints generated automatically from user feedbacks and sample documents. The quality of local metric is improved by incorporating more precise constraints. Improving the quality of local metric will in return enhance the clustering performance. We have conducted extensive experiments on real-world news documents. The results demonstrate that user feedback information coupled with local metric learning can dramatically improve the clustering performance.

This paper is substantially supported by grants from the Research Grant Council of the Hong Kong Special Administrative Region, China (Project Nos: CUHK 4179/03E and CUHK4193/04E), the Direct Grant of the Faculty of Engineering, CUHK (Project Code: 2050363), and CUHK Strategic Grant (No: 4410001). This work is also affiliated with the Microsoft-CUHK Joint Laboratory for Human-centric Computing and Interface Technologies.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 84.99; Price excludes VAT (USA)

Softcover Book: USD 109.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

Preview

Unable to display preview. Download preview PDF.

References

Basu, S., Banerjee, A., Mooney, R.: Semi-supervised clustering by seeding. In: Proceedings of the Nineteenth International conference on Machine Learning (2002)
Google Scholar
Basu, S., Bilenko, M., Mooney, R.J.: A probabilistic framework for semisupervised clustering. In: Proceedings of the tenth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 59–68 (2004)
Google Scholar
Bilenko, M., Basu, S., Mooney, R.J.: Integrating constraints and metric learning in semi-supervised clustering. In: Proceedings of the International Conference on Machine Learning (2004)
Google Scholar
Demiriz, A., Bennett, K., Embrechts, M.: Semi-supervised clustering using genetic algorithms. In: Artificial Neural Networks In Engineering (1999)
Google Scholar
Dhillon, I.S., Modha, D.S.: Concept decompositions for large sparse text data using clustering. Machine Learning 42(1), 143–175 (2001)
Article MATH Google Scholar
Frigui, H., Nasraoui, O.: Unsupervised learning of prototypes and attribute weights. Pattern recognition 37(3), 567–581 (2004)
Article Google Scholar
Jing, L., Ng, M.K., Xu, J., Huang, J.Z.: Subspace clustering of text documents with feature weighting K-Means algorithm. In: Proceedings of the Pacific-Asia Conference on Knowledge Discovery and Data Mining, pp. 802–812 (2005)
Google Scholar
Rose, T., Stevenson, M., Whitehea, M.: The Reuters Corpus Volume 1– from yesterday’s news to tomorrow’s language resources. In: Proceedings of Third International Conference on Language Resources and Evaluation (2002)
Google Scholar
Wagstaff, K., Cardie, C.: Clustering with instance-level constraints. In: Proceedings of the Seventeenth International Conference on Machine Learning (2000)
Google Scholar
Xing, E.P., Ng, A.Y., Jordan, M., Russell, S.: Distance metric learning, with application to clustering with side-information. Advances in NIPS 15 (2003)
Google Scholar
Zhao, Y., Karypis, G.: Hierarchical clustering algorithms for document datasets. Data Mining and Knowledge Discovery 10, 141–168 (2005)
Article MathSciNet Google Scholar

Download references

Author information

Authors and Affiliations

Department of Systems Engineering and Engineering Management, The Chinese University of Hong Kong, Shatin, Hong Kong
Ruizhang Huang, Zhigang Zhang & Wai Lam

Authors

Ruizhang Huang
View author publications
You can also search for this author in PubMed Google Scholar
Zhigang Zhang
View author publications
You can also search for this author in PubMed Google Scholar
Wai Lam
View author publications
You can also search for this author in PubMed Google Scholar

Editor information

Editors and Affiliations

Department of Computer Science, National University of Singapore, 3 Science Drive 2, 117543, Singapore
Hwee Tou Ng
Institute for Infocomm Research, 21 Heng Mui Keng Terrace, 119613, Singapore
Mun-Kew Leong
Department of Computer Science, School of Computing, National University of Singapore, 117543, Singapore
Min-Yen Kan
Institute for Infocomm Research, 21 Heng Mui Keng Terrace, P.O. Box, 119613, Singapore
Donghong Ji

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Huang, R., Zhang, Z., Lam, W. (2006). Text Clustering with Limited User Feedback Under Local Metric Learning. In: Ng, H.T., Leong, MK., Kan, MY., Ji, D. (eds) Information Retrieval Technology. AIRS 2006. Lecture Notes in Computer Science, vol 4182. Springer, Berlin, Heidelberg. https://doi.org/10.1007/11880592_11

Download citation

DOI: https://doi.org/10.1007/11880592_11
Publisher Name: Springer, Berlin, Heidelberg
Print ISBN: 978-3-540-45780-0
Online ISBN: 978-3-540-46237-8
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics