Deep Learning for Binaural Sound Source Localization with Low Signal-to-noise Ratio
Fengnian Zhao, Ruwei Li, Dongmei Pan
Abstract
Open-access reader
Fengnian Zhao, Ruwei Li, Dongmei Pan
Abstract
Open-access reader
Abstract A novel deep learning (DL) method is proposed for binaural sound source localization with low SNR. Firstly, the binaural sound signals are decomposed into several channels by using Gammatone filter. Secondly, the 4 feature parameters of Head-related Transfer Function, interaural time difference (ITD), interaural coherence (IC), interaural level difference (ILD), and interaural phase difference (IPD) are extracted. Thirdly, ITD and IC go through a Deep Belief Network (DBN) to determine the quadrant of the sound source and reduce the positioning range. Then, ITD, IC, ILD, and IPD go through a Deep Neural Network (DNN) to obtain the azimuthal angle within 90 degrees. Experimental results show that the proposed algorithm can solve the front-back confusion, and obtain a superior performance with lower complexity and higher precision under low SNR conditions.
OpenAlex reports 5 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Abstract A novel deep learning (DL) method is proposed for binaural sound source localization with low SNR. Firstly, the binaural sound signals are decomposed into several channels by using Gammatone filter. Secondly, the 4 feature parameters of Head-related Transfer Function, interaural time difference (ITD), interaural coherence (IC), interaural level difference (ILD), and interaural phase difference (IPD) are extracted. Thirdly, ITD and IC go through a Deep Belief Network (DBN) to determine the quadrant of the sound source and reduce the positioning range. Then, ITD, IC, ILD, and IPD go through a Deep Neural Network (DNN) to obtain the azimuthal angle within 90 degrees. Experimental results show that the proposed algorithm can solve the front-back confusion, and obtain a superior performance with lower complexity and higher precision under low SNR conditions.
Key concepts: Interaural time difference, Binaural recording, Sound localization, Head-related transfer function, Computer science, Azimuth, Acoustics, Speech recognition