Influence of Autocorrelation Lag Ranges on Robust Speech Recognition
Benjamin J. Shannon, Kuldip K. Paliwal
Abstract
Benjamin J. Shannon, Kuldip K. Paliwal
Abstract
It is generally believed that the lower-lag autocorrelation coefficients carry information about the spectral envelope and the higher-lag autocorrelation coefficients are more related to pitch information. In this paper, we use lower-lag and higher-lag ranges of the autocorrelation function separately for deriving speech recognition features, and investigate their role in terms of speech recognition performance. The state-of-the-art MFCC (mel frequency cepstral coefficient) features use the whole autocorrelation function in their computation and are used here as a benchmark in our experiments. Our recognition results from the Aurora II corpus show that the higher-lag autocorrelation coefficients perform as well as the whole autocorrelation function for clean speech, and provide better performance for noisy speech, while lower-lag autocorrelation coefficients are not as effective in this aspect.
OpenAlex reports 7 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
It is generally believed that the lower-lag autocorrelation coefficients carry information about the spectral envelope and the higher-lag autocorrelation coefficients are more related to pitch information. In this paper, we use lower-lag and higher-lag ranges of the autocorrelation function separately for deriving speech recognition features, and investigate their role in terms of speech recognition performance. The state-of-the-art MFCC (mel frequency cepstral coefficient) features use the whole autocorrelation function in their computation and are used here as a benchmark in our experiments. Our recognition results from the Aurora II corpus show that the higher-lag autocorrelation coefficients perform as well as the whole autocorrelation function for clean speech, and provide better performance for noisy speech, while lower-lag autocorrelation coefficients are not as effective in this aspect.
Key concepts: Autocorrelation, Lag, Autocorrelation technique, Mel-frequency cepstrum, Speech recognition, Spectral envelope, Computer science, Benchmark (surveying)