2024Unpublished venueRequires access

PESQ enhancement for decoded speech audio signals using complex convolutional recurrent neural network

Farah Shahhoud, Ahmad Ali Deeb, Valery I. Terekhov

Open publisher page 4 citations

Abstract

In telecommunication, speech signals need to be coded to ensure good transmission quality of the signals to far receivers. The objective of the codec system is to compress the signals to as few bits as possible while keeping good quality of the output signals. Unfortunately, encoding speech signals decreases their quality, especially when there are limits on the sent bitrates, which cause uncomfortable experience of communication for both sides. In this paper, a speech enhancement filter is proposed, which reduces distortions of the decoded speech at low bitrates to make it more pleasant, reduce listening effort and improve the intelligibility of the signal. The filter is a complex convolutional neural network model with Long-Short Term Memory (LSTM) units, that enhances the speech signals decoded by Enhanced Voice Services (EVS) codec. Our model was trained using perceptual metric for speech quality evaluation (PMSQE) which is a more intelligent loss function. In result, the Perceptual Evaluation of Speech Quality (PESQ) score evaluation of speech signals, enhanced by the model, has a raise by 0.2 on average for different bitrates signals.

About this research paper

What this paper is about

In telecommunication, speech signals need to be coded to ensure good transmission quality of the signals to far receivers. The objective of the codec system is to compress the signals to as few bits as possible while keeping good quality of the output signals. Unfortunately, encoding speech signals decreases their quality, especially when there are limits on the sent bitrates, which cause uncomfortable experience of communication for both sides. In this paper, a speech enhancement filter is proposed, which reduces distortions of the decoded speech at low bitrates to make it more pleasant, reduce listening effort and improve the intelligibility of the signal. The filter is a complex convolutional neural network model with Long-Short Term Memory (LSTM) units, that enhances the speech signals decoded by Enhanced Voice Services (EVS) codec. Our model was trained using perceptual metric for speech quality evaluation (PMSQE) which is a more intelligent loss function. In result, the Perceptual Evaluation of Speech Quality (PESQ) score evaluation of speech signals, enhanced by the model, has a raise by 0.2 on average for different bitrates signals.

Why it matters

OpenAlex reports 4 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

In telecommunication, speech signals need to be coded to ensure good transmission quality of the signals to far receivers. The objective of the codec system is to compress the signals to as few bits as possible while keeping good quality of the output signals. Unfortunately, encoding speech signals decreases their quality, especially when there are limits on the sent bitrates, which cause uncomfortable experience of communication for both sides. In this paper, a speech enhancement filter is proposed, which reduces distortions of the decoded speech at low bitrates to make it more pleasant, reduce listening effort and improve the intelligibility of the signal. The filter is a complex convolutional neural network model with Long-Short Term Memory (LSTM) units, that enhances the speech signals decoded by Enhanced Voice Services (EVS) codec. Our model was trained using perceptual metric for speech quality evaluation (PMSQE) which is a more intelligent loss function. In result, the Perceptual Evaluation of Speech Quality (PESQ) score evaluation of speech signals, enhanced by the model, has a raise by 0.2 on average for different bitrates signals.

Key concepts: PESQ, Computer science, Speech recognition, Speech coding, Codec2, Codec, PSQM, Voice activity detection

Related papers

Back to paper searchBrowse research topicsOriginal source
PESQ enhancement for decoded speech audio signals using complex convolutional recurrent neural network — Research Paper | ScholarLens