Finding People Frequently Appearing in News
We propose a graph based method to improve the performance of person queries in large news video collections. The method benefits from the multi-modal structure of videos and integrates text and face information. Using the idea that a person appears more frequently when his/her name is mentioned, we first use the speech transcript text to limit our search space for a query name. Then, we construct a similarity graph with nodes corresponding to all of the faces in the search space, and the edges corresponding to similarity of the faces. With the assumption that the images of the query name will be more similar to each other than to other images, the problem is then transformed into finding the densest component in the graph corresponding to the images of the query name. The same graph algorithm is applied for detecting and removing the faces of the anchorpeople in an unsupervised way. The experiments are conducted on 229 news videos provided by NIST for TRECVID 2004. The results show that proposed method outperforms the text only based methods and provides cues for recognition of faces on the large scale.
KeywordsTurkey Acoustics Boris
Unable to display preview. Download preview PDF.
- 1.Trec video retrieval evaluation (2004), http://www-nlpir.nist.gov/projects/trecvid/
- 2.Gross, R., Baker, S., Matthews, I., Kanade, T.: Face recognition across pose and illumination. In: Li, S.Z., Jain, A.K. (eds.) Handbook of Face Recognition. Springer, Heidelberg (2004)Google Scholar
- 4.Satoh, S., Kanade, T.: Name-it: Association of face and name in video. In: Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (1997)Google Scholar
- 5.Berg, T., Berg, A.C., Edwards, J., Forsyth, D.: Who is in the picture. In: Neural Information Processing Systems (NIPS) (2004)Google Scholar
- 6.Chen, M.Y., Hauptmann, A.: Searching for a specific person in broadcast news video. In: International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2004), Montreal, Canada (2004)Google Scholar
- 10.Berg, T., Berg, A.C., Edwards, J., Maire, M., White, R., Teh, Y.W., Learned-Miller, E., Forsyth, D.: Faces and names in the news. In: IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) (2004)Google Scholar
- 11.Ozkan, D., Duygulu, P.: Interesting faces in the news. In: IEEE Conf. on Computer Vision and Pattern Recognition (to appear, 2006)Google Scholar
- 12.Gauvain, J., Lamel, L., Adda, G.: The limsi broadcast news transcription system. Speech Communication 37(1-2) (2002)Google Scholar
- 13.Mikolajczyk, K.: Face detector. INRIA Rhone-Alpes, Ph.D Report (2004)Google Scholar
- 14.Lowe, D.G.: Distinctive image features from scale-invariant keypoints. International Journal of Computer Vision 60(2) (2004)Google Scholar