2012•Signal ProcessingRequires access

Voice Conversion Using Spectrum with Super-Segment Prosody Features

YU Yi-biao

Open publisher page 0 citations

Abstract

Only few of prosody features such as fundamental frequency is used in common voice conversion system,so the conversion speech has weak target tendency and poor quality especially when speakers have strong speaking styles.In this paper,a new conversion method based on short-time spectrum and prosodies such as pitch contour,duration,pause and stress is proposed.Pitch contour is first described by pitch target model and then trained by Gaussian mixture model (GMM ),the other prosodies are modeled by single Gaussian distribution model after statistical analysis.The experiment result show the target tendentiousness and naturalness of converted speech are well improved after use of rich prosody features comparing with traditional system.

About this research paper

What this paper is about

Only few of prosody features such as fundamental frequency is used in common voice conversion system,so the conversion speech has weak target tendency and poor quality especially when speakers have strong speaking styles.In this paper,a new conversion method based on short-time spectrum and prosodies such as pitch contour,duration,pause and stress is proposed.Pitch contour is first described by pitch target model and then trained by Gaussian mixture model (GMM ),the other prosodies are modeled by single Gaussian distribution model after statistical analysis.The experiment result show the target tendentiousness and naturalness of converted speech are well improved after use of rich prosody features comparing with traditional system.

Why it matters

A significance statement is not available in the OpenAlex record.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Only few of prosody features such as fundamental frequency is used in common voice conversion system,so the conversion speech has weak target tendency and poor quality especially when speakers have strong speaking styles.In this paper,a new conversion method based on short-time spectrum and prosodies such as pitch contour,duration,pause and stress is proposed.Pitch contour is first described by pitch target model and then trained by Gaussian mixture model (GMM ),the other prosodies are modeled by single Gaussian distribution model after statistical analysis.The experiment result show the target tendentiousness and naturalness of converted speech are well improved after use of rich prosody features comparing with traditional system.

Key concepts: Naturalness, Prosody, Pitch contour, Speech recognition, Gaussian, Mixture model, Computer science, Speech synthesis

Related papers

Back to paper searchBrowse research topicsOriginal source
Voice Conversion Using Spectrum with Super-Segment Prosody Features — Research Paper | ScholarLens