Low‐Bit‐Rate Speech Coding
Miguel Arjona Ramírez, Màrio Minami
Abstract
Miguel Arjona Ramírez, Màrio Minami
Abstract
Abstract This article is focused on speech coding methods for achieving communication quality speech at bit rates of 4 kbit/s and lower. The speech coding techniques are based on an all‐pole model of the vocal tract which may be implemented in the time domain with appropriately selected excitation functions or else may be fit to a spectral analysis of the speech signal. Three main types of coders are described below. Code‐excited linear prediction (CELP) coders select their excitation from waveform codebooks using analysis‐by‐synthesis closed‐loop techniques, which need to be supplemented by speech classification and open‐loop parametric techniques for keeping up with quality at lower rates. The prototypical sinusoidal coder (SC) has a bank of oscillators for signal synthesis, driven by a model of the magnitude spectrum. However, phase regeneration is important in enhancing speech reconstruction at low rates. Waveform interpolation (WI) coders afford a wider time‐frequency footprint for the representation of the excitation, showing a good potential for achieving toll quality at bit rates below 4 kbit/s.
OpenAlex reports 6 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Abstract This article is focused on speech coding methods for achieving communication quality speech at bit rates of 4 kbit/s and lower. The speech coding techniques are based on an all‐pole model of the vocal tract which may be implemented in the time domain with appropriately selected excitation functions or else may be fit to a spectral analysis of the speech signal. Three main types of coders are described below. Code‐excited linear prediction (CELP) coders select their excitation from waveform codebooks using analysis‐by‐synthesis closed‐loop techniques, which need to be supplemented by speech classification and open‐loop parametric techniques for keeping up with quality at lower rates. The prototypical sinusoidal coder (SC) has a bank of oscillators for signal synthesis, driven by a model of the magnitude spectrum. However, phase regeneration is important in enhancing speech reconstruction at low rates. Waveform interpolation (WI) coders afford a wider time‐frequency footprint for the representation of the excitation, showing a good potential for achieving toll quality at bit rates below 4 kbit/s.
Key concepts: Code-excited linear prediction, Speech coding, Linear predictive coding, Speech recognition, Computer science, Harmonic Vector Excitation Coding, Codec2, Waveform