HMM adaptation and voice conversion for the synthesis of child speech: a comparison
Oliver Watts, Junichi Yamagishi, Simon King, Kay Berkling
Abstract
Open-access reader
Oliver Watts, Junichi Yamagishi, Simon King, Kay Berkling
Abstract
Open-access reader
This study compares two different methodologies for producing data-driven synthesis of child speech from existing systems that have been trained on the speech of adults. On one hand, an existing statistical parametric synthesiser is transformed using model adaptation techniques, informed by linguistic and prosodic knowledge, to the speaker characteristics of a child speaker. This is compared with the application of voice conversion techniques to convert the output of an existing waveform concatenation synthesiser with no explicit linguistic or prosodic knowledge. In a subjective evaluation of the similarity of synthetic speech to natural speech from the target speaker, the HMM-based systems evaluated are generally preferred, although this is at least in part due to the higher dimensional acoustic features supported by these techniques. Index Terms: child speech, statistical parametric speech synthesis, HMM-based speech synthesis, voice conversion, HTS, Average Voice Models, Festival
OpenAlex reports 5 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
This study compares two different methodologies for producing data-driven synthesis of child speech from existing systems that have been trained on the speech of adults. On one hand, an existing statistical parametric synthesiser is transformed using model adaptation techniques, informed by linguistic and prosodic knowledge, to the speaker characteristics of a child speaker. This is compared with the application of voice conversion techniques to convert the output of an existing waveform concatenation synthesiser with no explicit linguistic or prosodic knowledge. In a subjective evaluation of the similarity of synthetic speech to natural speech from the target speaker, the HMM-based systems evaluated are generally preferred, although this is at least in part due to the higher dimensional acoustic features supported by these techniques. Index Terms: child speech, statistical parametric speech synthesis, HMM-based speech synthesis, voice conversion, HTS, Average Voice Models, Festival
Key concepts: Speech synthesis, Speech recognition, Concatenation (mathematics), Computer science, Hidden Markov model, Parametric statistics, Voice analysis, Adaptation (eye)