The effect of phase information in speech enhancement and speech recognition
Mahsa Sadat Elyasi Langarani, Hadi Veisi, Hossein Sameti
Abstract
Mahsa Sadat Elyasi Langarani, Hadi Veisi, Hossein Sameti
Abstract
The majority of speech enhancement methods perform noise removal in spectral domain and construct the enhanced speech signal from the estimated magnitude of clean speech and the phase of the noisy speech. In this paper, we show that by incorporating the phase information in the enhancement process, the quality and intelligibility of speech signal are improved. In our investigations, the minimum mean-square error short-time spectral amplitude and MMSE log-spectral amplitude methods are used to estimate the magnitude spectrum of speech signal. By conducting six classes of experiments, it is shown that by taking the phase information into account, overall SNR and PESQ measures are improved. In other experiments, the results of speech enhancement system are fed into a speech recognition system and it is shown that by incorporating accurate phase information, the intelligibility of the speech signal is improved.
OpenAlex reports 10 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
The majority of speech enhancement methods perform noise removal in spectral domain and construct the enhanced speech signal from the estimated magnitude of clean speech and the phase of the noisy speech. In this paper, we show that by incorporating the phase information in the enhancement process, the quality and intelligibility of speech signal are improved. In our investigations, the minimum mean-square error short-time spectral amplitude and MMSE log-spectral amplitude methods are used to estimate the magnitude spectrum of speech signal. By conducting six classes of experiments, it is shown that by taking the phase information into account, overall SNR and PESQ measures are improved. In other experiments, the results of speech enhancement system are fed into a speech recognition system and it is shown that by incorporating accurate phase information, the intelligibility of the speech signal is improved.
Key concepts: PESQ, Intelligibility (philosophy), Speech enhancement, Speech recognition, Computer science, Voice activity detection, Linear predictive coding, Speech processing