Synthesis of continuous speech by concatenation of isolated words
Mark A. Randolph, Victor W. Zue
Abstract
Mark A. Randolph, Victor W. Zue
Abstract
This paper reports a feasibility study of synthesizing continuous speech by concatentation of modified isolated word templates. Continuous speech obtained simply by word template concatenation has several inherent problems. It cannot account for the prosodic features normally found in continuous speech, nor can it account for coarticulation; the natural transitions that occur at word boundaries. In an effort to alleviate the first problems, the synthesizer is provided with information describing the duration, the fundamental frequency (f0) contour, and the energy contour of each sentence. The synthesizer draws from a dictionary of word templates each stored as a sequence of LPC parameters. Before concatenation, the time scales of the synthesis templates are nonlinearly warped using the alignment path obtained from a level building connected speech recognition algorithm. In addition, the f0 and energy contours of the word templates are modified. To account for coarticulation, the various parameters are smoothed at word boundaries. The goal of this synthesis system is an output information rate of 200 bits per second of speech. A demonstration tape will be played.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
This paper reports a feasibility study of synthesizing continuous speech by concatentation of modified isolated word templates. Continuous speech obtained simply by word template concatenation has several inherent problems. It cannot account for the prosodic features normally found in continuous speech, nor can it account for coarticulation; the natural transitions that occur at word boundaries. In an effort to alleviate the first problems, the synthesizer is provided with information describing the duration, the fundamental frequency (f0) contour, and the energy contour of each sentence. The synthesizer draws from a dictionary of word templates each stored as a sequence of LPC parameters. Before concatenation, the time scales of the synthesis templates are nonlinearly warped using the alignment path obtained from a level building connected speech recognition algorithm. In addition, the f0 and energy contours of the word templates are modified. To account for coarticulation, the various parameters are smoothed at word boundaries. The goal of this synthesis system is an output information rate of 200 bits per second of speech. A demonstration tape will be played.
Key concepts: Coarticulation, Concatenation (mathematics), Computer science, Speech synthesis, Word (group theory), Template, Speech recognition, Sentence