Weak speech recovery for single-channel speech enhancement
Arthur Wong Kok Ming, Siow Yong Low
Abstract
Arthur Wong Kok Ming, Siow Yong Low
Abstract
Numerous speech enhancement strategies have been proposed to counteract the background noise present in communication and thus improve speech quality. However, such an improvement does not imply an improvement in speech intelligibility. Therefore, the issue become challenging especially it concerns single-channel situations. In this study, we present an approach to improve both speech quality and intelligibility, that is to recover weak speech. Here, weak speech is determined based on the pitch estimation and signal to noise ratio (SNR). With both parameters, we review the methodology of spectral subtraction and pitch estimation and then explore the possibility of improving the speech by incorporating the pitch information.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Numerous speech enhancement strategies have been proposed to counteract the background noise present in communication and thus improve speech quality. However, such an improvement does not imply an improvement in speech intelligibility. Therefore, the issue become challenging especially it concerns single-channel situations. In this study, we present an approach to improve both speech quality and intelligibility, that is to recover weak speech. Here, weak speech is determined based on the pitch estimation and signal to noise ratio (SNR). With both parameters, we review the methodology of spectral subtraction and pitch estimation and then explore the possibility of improving the speech by incorporating the pitch information.
Key concepts: Intelligibility (philosophy), Speech enhancement, Speech recognition, Computer science, Voice activity detection, Speech processing, Linear predictive coding, PSQM