2021•2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)Requires access

Vibrato Learning in Multi-Singer Singing Voice Synthesis

Ruolan Liu, Xue Qin Wen, Lu Chunhui, Liming Song, June Sig Sung

Open publisher page 9 citations

Abstract

Decent vibratos are a trait of good vocal training, often associated with perceived level of singing skill. In this paper we present a system for multi-singer singing voice synthesis, which is capable of producing high quality singing with convincing, controllable vibratos, and can also synthesize natural singing voices for target speakers with only speech data. This is enabled by using a unified speech-and-singing acoustic model that not only bridges the modality gap but also helps make best use of both types of data. The acoustic model exposes the full F0 contour therefore allowing explicitly modelling of F0 characteristics specific to singing voice. We observe that short-time Fourier transform of the F0 contour sparsely encode vibrato characteristics, and derive a learning objective therefrom for improved vibrato production. Control of the synthesized vibrato extent is possible by wiring in a supervised “extent” neuron and expose it to outer system. Experimental results confirm the effectiveness of proposed objective in producing good vibratos and improving overall perceived singing voice quality.

About this research paper

What this paper is about

Decent vibratos are a trait of good vocal training, often associated with perceived level of singing skill. In this paper we present a system for multi-singer singing voice synthesis, which is capable of producing high quality singing with convincing, controllable vibratos, and can also synthesize natural singing voices for target speakers with only speech data. This is enabled by using a unified speech-and-singing acoustic model that not only bridges the modality gap but also helps make best use of both types of data. The acoustic model exposes the full F0 contour therefore allowing explicitly modelling of F0 characteristics specific to singing voice. We observe that short-time Fourier transform of the F0 contour sparsely encode vibrato characteristics, and derive a learning objective therefrom for improved vibrato production. Control of the synthesized vibrato extent is possible by wiring in a supervised “extent” neuron and expose it to outer system. Experimental results confirm the effectiveness of proposed objective in producing good vibratos and improving overall perceived singing voice quality.

Why it matters

OpenAlex reports 9 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Decent vibratos are a trait of good vocal training, often associated with perceived level of singing skill. In this paper we present a system for multi-singer singing voice synthesis, which is capable of producing high quality singing with convincing, controllable vibratos, and can also synthesize natural singing voices for target speakers with only speech data. This is enabled by using a unified speech-and-singing acoustic model that not only bridges the modality gap but also helps make best use of both types of data. The acoustic model exposes the full F0 contour therefore allowing explicitly modelling of F0 characteristics specific to singing voice. We observe that short-time Fourier transform of the F0 contour sparsely encode vibrato characteristics, and derive a learning objective therefrom for improved vibrato production. Control of the synthesized vibrato extent is possible by wiring in a supervised “extent” neuron and expose it to outer system. Experimental results confirm the effectiveness of proposed objective in producing good vibratos and improving overall perceived singing voice quality.

Key concepts: Vibrato, Singing, Speech recognition, Computer science, Quality (philosophy), Natural (archaeology), Acoustics, Philosophy

Related papers

Back to paper searchBrowse research topicsOriginal source
Vibrato Learning in Multi-Singer Singing Voice Synthesis — Research Paper | ScholarLens