A noise robust speech activity detection algorithm
Bala Harsha
Abstract
Bala Harsha
Abstract
This paper proposes an efficient and robust speech starting and end point detection method, which improve the performance for speech recognition in a noisy environment. The proposed method designs a series of speech/non-speech classifiers for voice activity detection and robust end-point detection using an 'adaptive thresholding' algorithm. The proposed method uses multiple features of speech for robust speech detection under noisy conditions, especially in an automotive environment. The key advantages of this method are its simple implementation and its low computational complexity. The proposed algorithm is used for isolated word recognition in a discontinuous speech recognition system. The performance of the proposed algorithm is measured in a simulated noisy environment with speech wave files recorded under noisy conditions.
OpenAlex reports 14 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
This paper proposes an efficient and robust speech starting and end point detection method, which improve the performance for speech recognition in a noisy environment. The proposed method designs a series of speech/non-speech classifiers for voice activity detection and robust end-point detection using an 'adaptive thresholding' algorithm. The proposed method uses multiple features of speech for robust speech detection under noisy conditions, especially in an automotive environment. The key advantages of this method are its simple implementation and its low computational complexity. The proposed algorithm is used for isolated word recognition in a discontinuous speech recognition system. The performance of the proposed algorithm is measured in a simulated noisy environment with speech wave files recorded under noisy conditions.
Key concepts: Computer science, Voice activity detection, Speech recognition, Thresholding, Noise (video), Speech enhancement, Word (group theory), Robustness (evolution)