PESQ enhancement for decoded speech audio signals using complex convolutional recurrent neural network
Farah Shahhoud, Ahmad Ali Deeb, Valery I. Terekhov
Abstract
Farah Shahhoud, Ahmad Ali Deeb, Valery I. Terekhov
Abstract
In telecommunication, speech signals need to be coded to ensure good transmission quality of the signals to far receivers. The objective of the codec system is to compress the signals to as few bits as possible while keeping good quality of the output signals. Unfortunately, encoding speech signals decreases their quality, especially when there are limits on the sent bitrates, which cause uncomfortable experience of communication for both sides. In this paper, a speech enhancement filter is proposed, which reduces distortions of the decoded speech at low bitrates to make it more pleasant, reduce listening effort and improve the intelligibility of the signal. The filter is a complex convolutional neural network model with Long-Short Term Memory (LSTM) units, that enhances the speech signals decoded by Enhanced Voice Services (EVS) codec. Our model was trained using perceptual metric for speech quality evaluation (PMSQE) which is a more intelligent loss function. In result, the Perceptual Evaluation of Speech Quality (PESQ) score evaluation of speech signals, enhanced by the model, has a raise by 0.2 on average for different bitrates signals.
OpenAlex reports 4 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
In telecommunication, speech signals need to be coded to ensure good transmission quality of the signals to far receivers. The objective of the codec system is to compress the signals to as few bits as possible while keeping good quality of the output signals. Unfortunately, encoding speech signals decreases their quality, especially when there are limits on the sent bitrates, which cause uncomfortable experience of communication for both sides. In this paper, a speech enhancement filter is proposed, which reduces distortions of the decoded speech at low bitrates to make it more pleasant, reduce listening effort and improve the intelligibility of the signal. The filter is a complex convolutional neural network model with Long-Short Term Memory (LSTM) units, that enhances the speech signals decoded by Enhanced Voice Services (EVS) codec. Our model was trained using perceptual metric for speech quality evaluation (PMSQE) which is a more intelligent loss function. In result, the Perceptual Evaluation of Speech Quality (PESQ) score evaluation of speech signals, enhanced by the model, has a raise by 0.2 on average for different bitrates signals.
Key concepts: PESQ, Computer science, Speech recognition, Speech coding, Codec2, Codec, PSQM, Voice activity detection