2013•Loughborough University Institutional Repository (Loughborough University)Open access

Speech separation with dereverberation-based pre-processing incorporating visual cues

Muhammad Salman Khan, Syed Mohsen Naqvi, Jonathon A. Chambers

Open full text 2 citations

Abstract

Humans are skilled in selectively extracting a single sound\nsource in the presence of multiple simultaneous sounds. They\n(individuals with normal hearing) can also robustly adapt to\nchanging acoustic environments with great ease. Need has\narisen to incorporate such abilities in machines which would\nenable multiple application areas such as human-computer\ninteraction, automatic speech recognition, hearing aids and\nhands-free telephony. This work addresses the problem of\nseparating multiple speech sources in realistic reverberant\nrooms using two microphones.\nDifferent monaural and binaural cues have previously\nbeen modeled in order to enable separation. Binaural spatial\ncues i.e. the interaural level difference (ILD) and the inter-\naural phase difference (IPD) have been modeled [1] in the\ntime-frequency (TF) domain that exploit the differences in\nthe intensity and the phase of the mixture signals (because of\nthe different spatial locations) observed by two microphones\n(or ears). The method performs well with no or little rever-\nberation but as the amount of reverberation increases and the\nsources approach each other, the binaural cues are distorted\nand the interaural cues become indistinct, hence, degrading\nthe separation performance. Thus, there is a demand for\nexploiting additional cues, and further signal processing is\nrequired at higher levels of reverberation.

Open-access reader

About this research paper

What this paper is about

Humans are skilled in selectively extracting a single sound\nsource in the presence of multiple simultaneous sounds. They\n(individuals with normal hearing) can also robustly adapt to\nchanging acoustic environments with great ease. Need has\narisen to incorporate such abilities in machines which would\nenable multiple application areas such as human-computer\ninteraction, automatic speech recognition, hearing aids and\nhands-free telephony. This work addresses the problem of\nseparating multiple speech sources in realistic reverberant\nrooms using two microphones.\nDifferent monaural and binaural cues have previously\nbeen modeled in order to enable separation. Binaural spatial\ncues i.e. the interaural level difference (ILD) and the inter-\naural phase difference (IPD) have been modeled [1] in the\ntime-frequency (TF) domain that exploit the differences in\nthe intensity and the phase of the mixture signals (because of\nthe different spatial locations) observed by two microphones\n(or ears). The method performs well with no or little rever-\nberation but as the amount of reverberation increases and the\nsources approach each other, the binaural cues are distorted\nand the interaural cues become indistinct, hence, degrading\nthe separation performance. Thus, there is a demand for\nexploiting additional cues, and further signal processing is\nrequired at higher levels of reverberation.

Why it matters

OpenAlex reports 2 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Humans are skilled in selectively extracting a single sound\nsource in the presence of multiple simultaneous sounds. They\n(individuals with normal hearing) can also robustly adapt to\nchanging acoustic environments with great ease. Need has\narisen to incorporate such abilities in machines which would\nenable multiple application areas such as human-computer\ninteraction, automatic speech recognition, hearing aids and\nhands-free telephony. This work addresses the problem of\nseparating multiple speech sources in realistic reverberant\nrooms using two microphones.\nDifferent monaural and binaural cues have previously\nbeen modeled in order to enable separation. Binaural spatial\ncues i.e. the interaural level difference (ILD) and the inter-\naural phase difference (IPD) have been modeled [1] in the\ntime-frequency (TF) domain that exploit the differences in\nthe intensity and the phase of the mixture signals (because of\nthe different spatial locations) observed by two microphones\n(or ears). The method performs well with no or little rever-\nberation but as the amount of reverberation increases and the\nsources approach each other, the binaural cues are distorted\nand the interaural cues become indistinct, hence, degrading\nthe separation performance. Thus, there is a demand for\nexploiting additional cues, and further signal processing is\nrequired at higher levels of reverberation.

Key concepts: Binaural recording, Monaural, Reverberation, Computer science, Speech recognition, Sound localization, Interaural time difference, Speech processing

Related papers

Back to paper searchBrowse research topicsOriginal source
Speech separation with dereverberation-based pre-processing incorporating visual cues — Research Paper | ScholarLens