A Robust Dual-Microphone Speech Source Localization Algorithm for Reverberant Environments
Yanmeng Guo, Xiaofei Wang, Chao Wu, Qiang Fu, Ning Ma, Guy J. Brown
Abstract
Yanmeng Guo, Xiaofei Wang, Chao Wu, Qiang Fu, Ning Ma, Guy J. Brown
Abstract
Speech source localization (SSL) using a microphone array \naims to estimate the direction-of-arrival (DOA) of the speech \nsource. However, its performance often degrades rapidly in reverberant \nenvironments. In this paper, a novel dual-microphone \nSSL algorithm is proposed to address this problem. First, the \ntime-frequency regions dominated by direct sound are extracted \nby tracking the envelopes of speech, reverberation and background \nnoise. The time-difference-of-arrival (TDOA) is then \nestimated by considering only these reliable regions. Second, \na bin-wise de-aliasing strategy is introduced to make better use \nof the DOA information carried at high frequencies, where the \nspatial resolution is higher and there is typically less corruption \nby diffuse noise. Our experiments show that when compared \nwith other widely-used algorithms, the proposed algorithm produces \nmore reliable performance in realistic reverberant environments.
OpenAlex reports 8 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Speech source localization (SSL) using a microphone array \naims to estimate the direction-of-arrival (DOA) of the speech \nsource. However, its performance often degrades rapidly in reverberant \nenvironments. In this paper, a novel dual-microphone \nSSL algorithm is proposed to address this problem. First, the \ntime-frequency regions dominated by direct sound are extracted \nby tracking the envelopes of speech, reverberation and background \nnoise. The time-difference-of-arrival (TDOA) is then \nestimated by considering only these reliable regions. Second, \na bin-wise de-aliasing strategy is introduced to make better use \nof the DOA information carried at high frequencies, where the \nspatial resolution is higher and there is typically less corruption \nby diffuse noise. Our experiments show that when compared \nwith other widely-used algorithms, the proposed algorithm produces \nmore reliable performance in realistic reverberant environments.
Key concepts: Reverberation, Multilateration, Computer science, Microphone, Aliasing, Microphone array, Speech recognition, Noise (video)