1982•The Journal of the Acoustical Society of AmericaOpen access

Synthesis of continuous speech by concatenation of isolated words

Mark A. Randolph, Victor W. Zue

Open full text 0 citations

Abstract

This paper reports a feasibility study of synthesizing continuous speech by concatentation of modified isolated word templates. Continuous speech obtained simply by word template concatenation has several inherent problems. It cannot account for the prosodic features normally found in continuous speech, nor can it account for coarticulation; the natural transitions that occur at word boundaries. In an effort to alleviate the first problems, the synthesizer is provided with information describing the duration, the fundamental frequency (f0) contour, and the energy contour of each sentence. The synthesizer draws from a dictionary of word templates each stored as a sequence of LPC parameters. Before concatenation, the time scales of the synthesis templates are nonlinearly warped using the alignment path obtained from a level building connected speech recognition algorithm. In addition, the f0 and energy contours of the word templates are modified. To account for coarticulation, the various parameters are smoothed at word boundaries. The goal of this synthesis system is an output information rate of 200 bits per second of speech. A demonstration tape will be played.

About this research paper

What this paper is about

This paper reports a feasibility study of synthesizing continuous speech by concatentation of modified isolated word templates. Continuous speech obtained simply by word template concatenation has several inherent problems. It cannot account for the prosodic features normally found in continuous speech, nor can it account for coarticulation; the natural transitions that occur at word boundaries. In an effort to alleviate the first problems, the synthesizer is provided with information describing the duration, the fundamental frequency (f0) contour, and the energy contour of each sentence. The synthesizer draws from a dictionary of word templates each stored as a sequence of LPC parameters. Before concatenation, the time scales of the synthesis templates are nonlinearly warped using the alignment path obtained from a level building connected speech recognition algorithm. In addition, the f0 and energy contours of the word templates are modified. To account for coarticulation, the various parameters are smoothed at word boundaries. The goal of this synthesis system is an output information rate of 200 bits per second of speech. A demonstration tape will be played.

Why it matters

A significance statement is not available in the OpenAlex record.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

This paper reports a feasibility study of synthesizing continuous speech by concatentation of modified isolated word templates. Continuous speech obtained simply by word template concatenation has several inherent problems. It cannot account for the prosodic features normally found in continuous speech, nor can it account for coarticulation; the natural transitions that occur at word boundaries. In an effort to alleviate the first problems, the synthesizer is provided with information describing the duration, the fundamental frequency (f0) contour, and the energy contour of each sentence. The synthesizer draws from a dictionary of word templates each stored as a sequence of LPC parameters. Before concatenation, the time scales of the synthesis templates are nonlinearly warped using the alignment path obtained from a level building connected speech recognition algorithm. In addition, the f0 and energy contours of the word templates are modified. To account for coarticulation, the various parameters are smoothed at word boundaries. The goal of this synthesis system is an output information rate of 200 bits per second of speech. A demonstration tape will be played.

Key concepts: Coarticulation, Concatenation (mathematics), Computer science, Speech synthesis, Word (group theory), Template, Speech recognition, Sentence

Related papers

Back to paper searchBrowse research topicsOriginal source
Synthesis of continuous speech by concatenation of isolated words — Research Paper | ScholarLens