Speech separation with dereverberation-based pre-processing incorporating visual cues
Muhammad Salman Khan, Syed Mohsen Naqvi, Jonathon A. Chambers
Abstract
Open-access reader
Muhammad Salman Khan, Syed Mohsen Naqvi, Jonathon A. Chambers
Abstract
Open-access reader
Humans are skilled in selectively extracting a single sound\nsource in the presence of multiple simultaneous sounds. They\n(individuals with normal hearing) can also robustly adapt to\nchanging acoustic environments with great ease. Need has\narisen to incorporate such abilities in machines which would\nenable multiple application areas such as human-computer\ninteraction, automatic speech recognition, hearing aids and\nhands-free telephony. This work addresses the problem of\nseparating multiple speech sources in realistic reverberant\nrooms using two microphones.\nDifferent monaural and binaural cues have previously\nbeen modeled in order to enable separation. Binaural spatial\ncues i.e. the interaural level difference (ILD) and the inter-\naural phase difference (IPD) have been modeled [1] in the\ntime-frequency (TF) domain that exploit the differences in\nthe intensity and the phase of the mixture signals (because of\nthe different spatial locations) observed by two microphones\n(or ears). The method performs well with no or little rever-\nberation but as the amount of reverberation increases and the\nsources approach each other, the binaural cues are distorted\nand the interaural cues become indistinct, hence, degrading\nthe separation performance. Thus, there is a demand for\nexploiting additional cues, and further signal processing is\nrequired at higher levels of reverberation.
OpenAlex reports 2 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Humans are skilled in selectively extracting a single sound\nsource in the presence of multiple simultaneous sounds. They\n(individuals with normal hearing) can also robustly adapt to\nchanging acoustic environments with great ease. Need has\narisen to incorporate such abilities in machines which would\nenable multiple application areas such as human-computer\ninteraction, automatic speech recognition, hearing aids and\nhands-free telephony. This work addresses the problem of\nseparating multiple speech sources in realistic reverberant\nrooms using two microphones.\nDifferent monaural and binaural cues have previously\nbeen modeled in order to enable separation. Binaural spatial\ncues i.e. the interaural level difference (ILD) and the inter-\naural phase difference (IPD) have been modeled [1] in the\ntime-frequency (TF) domain that exploit the differences in\nthe intensity and the phase of the mixture signals (because of\nthe different spatial locations) observed by two microphones\n(or ears). The method performs well with no or little rever-\nberation but as the amount of reverberation increases and the\nsources approach each other, the binaural cues are distorted\nand the interaural cues become indistinct, hence, degrading\nthe separation performance. Thus, there is a demand for\nexploiting additional cues, and further signal processing is\nrequired at higher levels of reverberation.
Key concepts: Binaural recording, Monaural, Reverberation, Computer science, Speech recognition, Sound localization, Interaural time difference, Speech processing