An Iterative Post-processing Approach for Speech Enhancement
Zhang Huimin, Xupeng Jia, Dongmei Li
Abstract
Zhang Huimin, Xupeng Jia, Dongmei Li
Abstract
Speech enhancement has been widely used in speech recognition, multimedia systems and hearing aids etc. In this study, we explore a new post-processing strategy for speech enhancement. The main goal of proposed post-processing method is to reduce speech distortion and improve speech quality and intelligibility after enhancement. First, a masking-based speech enhancement system based on deep neural network is implemented. Then, an iterative global variance equalization post-processing is proposed to adopt on estimated masks. We evaluate the intelligibility and quality of enhanced speech and observe that the proposed post-processing method achieves higher speech intelligibility and less speech distortion at low signal-to-noise ratios (SNRs) comparing to the baseline system without post-processing or previous post-processing methods. The experiments under unseen noises also show that the proposed post-processing strategy can improve the model generalization at multiple noise types.
OpenAlex reports 2 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Speech enhancement has been widely used in speech recognition, multimedia systems and hearing aids etc. In this study, we explore a new post-processing strategy for speech enhancement. The main goal of proposed post-processing method is to reduce speech distortion and improve speech quality and intelligibility after enhancement. First, a masking-based speech enhancement system based on deep neural network is implemented. Then, an iterative global variance equalization post-processing is proposed to adopt on estimated masks. We evaluate the intelligibility and quality of enhanced speech and observe that the proposed post-processing method achieves higher speech intelligibility and less speech distortion at low signal-to-noise ratios (SNRs) comparing to the baseline system without post-processing or previous post-processing methods. The experiments under unseen noises also show that the proposed post-processing strategy can improve the model generalization at multiple noise types.
Key concepts: Speech enhancement, Computer science, Speech recognition, Intelligibility (philosophy), Speech processing, Voice activity detection, Signal processing, PSQM