Sound Recognition in Mixtures

  • Juhan Nam
  • Gautham J. Mysore
  • Paris Smaragdis
Part of the Lecture Notes in Computer Science book series (LNCS, volume 7191)

Abstract

In this paper, we describe a method for recognizing sound sources in a mixture. While many audio-based content analysis methods focus on detecting or classifying target sounds in a discriminative manner, we approach this as a regression problem, in which we estimate the relative proportions of sound sources in the given mixture. Using source separation ideas based on probabilistic latent component analysis, we directly estimate these proportions from the mixture without actually separating the sources. We also introduce a method for learning a transition matrix to temporally constrain the problem. We demonstrate the proposed method on a mixture of five classes of sounds and show that it is quite effective in correctly estimating the relative proportions of the sounds in the mixture.

Preview

Unable to display preview. Download preview PDF.

Unable to display preview. Download preview PDF.

References

  1. 1.
    Radhakrishnan, R., Xiong, Z., Otsuka, I.: A Content-Adaptive Analysis and Representation Framework for Audio Event Discovery from Unscripted Multimedia. EURASIP Journal on Applied Signal Processing, 1–24 (2006)Google Scholar
  2. 2.
    Li, Y., Dorai, C.: Instructional Video Content Analysis Using Audio Information. IEEE TASLP 14(6) (2006)Google Scholar
  3. 3.
    Tran, H.D., Li, H.: Sound Event Recognition With Probabilistic Distance SVMs. IEEE TASLP 19(6) (2011)Google Scholar
  4. 4.
    Smaragdis, P., Raj, B., Shashanka, M.: A probabilistic latent variable model for acoustic modeling. In: Advances in Models for Acoustic Processing, NIPS (2006)Google Scholar
  5. 5.
    Smaragdis, P., Raj, B., Shashanka, M.: Supervised and Semi-Supervised Separation of Sounds from Single-Channel Mixtures. In: Davies, M.E., James, C.J., Abdallah, S.A., Plumbley, M.D. (eds.) ICA 2007. LNCS, vol. 4666, pp. 414–421. Springer, Heidelberg (2007)CrossRefGoogle Scholar

Copyright information

© Springer-Verlag Berlin Heidelberg 2012

Authors and Affiliations

  • Juhan Nam
    • 1
  • Gautham J. Mysore
    • 2
  • Paris Smaragdis
    • 2
    • 3
  1. 1.Center for Computer Research in Music and AcousticsStanford UniversityUSA
  2. 2.Advanced Technology LabsAdobe Systems Inc.USA
  3. 3.University of Illinois at Urbana-ChampaignUSA

Personalised recommendations