A new feature analysis method for robust ASR in reverberant environments based on the harmonic structure of speech
Rico Petrick, Kevin Lohde, Mike Lorenz, Ruediger Hoffmann
Abstract
Rico Petrick, Kevin Lohde, Mike Lorenz, Ruediger Hoffmann
Abstract
This article proposes a new signal analysis method for automatic speech recognition designed to aim high robustness against distortions caused by room reverberation. The method is initially named Harmonicity based Feature Analysis (HFA) and implements the following three ideas: (i) reconstruction of a spectrum from the harmonic components (assumed to be undistorted) of a voiced speech spectrum. (ii) suppression of disturbing reverberation in unvoiced spectra coming from previous voiced sections. (iii) high frequency regions are not affected by HFA since they have negligible effect on the recognition rate. HFA works on the basis of fundamental frequency estimation and voiced/unvoiced decision. Evaluation results show significant improvement of the recognition performance over a wide range of reverberant conditions while using HFA in connection with reverberant training. Apart from good performance, advantages of HFA compared to state of the art dereverberation approaches are real time processing (no adaptation time) and robustness against changes of the room impulse response.
OpenAlex reports 6 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
This article proposes a new signal analysis method for automatic speech recognition designed to aim high robustness against distortions caused by room reverberation. The method is initially named Harmonicity based Feature Analysis (HFA) and implements the following three ideas: (i) reconstruction of a spectrum from the harmonic components (assumed to be undistorted) of a voiced speech spectrum. (ii) suppression of disturbing reverberation in unvoiced spectra coming from previous voiced sections. (iii) high frequency regions are not affected by HFA since they have negligible effect on the recognition rate. HFA works on the basis of fundamental frequency estimation and voiced/unvoiced decision. Evaluation results show significant improvement of the recognition performance over a wide range of reverberant conditions while using HFA in connection with reverberant training. Apart from good performance, advantages of HFA compared to state of the art dereverberation approaches are real time processing (no adaptation time) and robustness against changes of the room impulse response.
Key concepts: Reverberation, Robustness (evolution), Speech recognition, Computer science, Impulse (physics), Speech processing, Impulse response, Feature extraction