A joint speech coding-enhancement algorithm for MBE vocoder
Huaiyu Dai, Zhigang Cao
Abstract
Huaiyu Dai, Zhigang Cao
Abstract
Multi-band excitation vocoder can yield high speech quality at medium and low rate speech coding, but its performance degrades dramatically in low SNR environments. A joint speech coding-enhancement algorithm is proposed which integrates MBE speech coding with speech enhancement. The original MBE coding algorithm is modified for better estimation of the MBE coding parameters in a noisy environment. Especially, the MMSE-LOG speech enhancement algorithm is incorporated in the MBE frequency domain calculation, and speech enhancement using the dual excitation speech model is adopted for post-processing of the synthesized speech. The objective measure and informal listening test indicate that significant improvement can be attained with an input SNR as low as -5 dB.
OpenAlex reports 2 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Multi-band excitation vocoder can yield high speech quality at medium and low rate speech coding, but its performance degrades dramatically in low SNR environments. A joint speech coding-enhancement algorithm is proposed which integrates MBE speech coding with speech enhancement. The original MBE coding algorithm is modified for better estimation of the MBE coding parameters in a noisy environment. Especially, the MMSE-LOG speech enhancement algorithm is incorporated in the MBE frequency domain calculation, and speech enhancement using the dual excitation speech model is adopted for post-processing of the synthesized speech. The objective measure and informal listening test indicate that significant improvement can be attained with an input SNR as low as -5 dB.
Key concepts: Computer science, Speech coding, Speech recognition, Codec2, Speech enhancement, Harmonic Vector Excitation Coding, Coding (social sciences), Linear predictive coding