Diphone synthesis using multipulse coding and a phase vecoder
Maja Stella, Francis Charpentier
Abstract
Maja Stella, Francis Charpentier
Abstract
Multipulse Linear Predictive Coding [1] has been shown to produce natural sounding speech at relatively low bit rates. So far, this technique has mostly been used for speech transmission or storage. In this paper, we show that a multipulse LPC synthesizer can also be used in a text-to-speech system based on diphone concatenation. The main problem is how to manipulate the prosodic parameters required for speech synthesis, and it is addressed here by a two-step procedure. First, a speech signal with relatively flat pitch contour is obtained by multipulse synthesis of concatenated diphones. Then the prosodic parameters of this signal are corrected using a special purpose phase vocoder. This method produces French synthetic speech of fairly good naturalness.
OpenAlex reports 6 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Multipulse Linear Predictive Coding [1] has been shown to produce natural sounding speech at relatively low bit rates. So far, this technique has mostly been used for speech transmission or storage. In this paper, we show that a multipulse LPC synthesizer can also be used in a text-to-speech system based on diphone concatenation. The main problem is how to manipulate the prosodic parameters required for speech synthesis, and it is addressed here by a two-step procedure. First, a speech signal with relatively flat pitch contour is obtained by multipulse synthesis of concatenated diphones. Then the prosodic parameters of this signal are corrected using a special purpose phase vocoder. This method produces French synthetic speech of fairly good naturalness.
Key concepts: Linear predictive coding, Computer science, Speech synthesis, Speech coding, Speech recognition, Concatenation (mathematics), Naturalness, Codec2