Some experiments on HMM speaker adaptation
A. Jarre, Roberto Pieraccini
Abstract
A. Jarre, Roberto Pieraccini
Abstract
The main problems with HMMs of sub-word units are the large amount of training data and computer time needed for estimating the parameters of the models. In some applications it is not proposable that a new speaker utters many hundreds of words to train the system, hence the interest arises for a quick adaptation based on some tens of training utterances. Two bounds are given for comparison with the results of the speaker adaptation, namely the recognition rates of speaker dependent and cross speaker recognition. Speaker dependent recognition is achieved by training the HMMs with nearly 1000 words uttered by the same speaker used in the tests. Cross speaker recognition, that gives a lower bound to the performance, concerns experiments in which the models were trained by a speaker different from that who uttered the test sentences. An adaptation algorithm using Parzen estimation and interpolation of the emission densities between the new and the old speaker models was investigated. It is able to give satisfactory recognition rates adapting the HMMs on the basis of only 40 training words uttered by the new speaker.
OpenAlex reports 3 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
The main problems with HMMs of sub-word units are the large amount of training data and computer time needed for estimating the parameters of the models. In some applications it is not proposable that a new speaker utters many hundreds of words to train the system, hence the interest arises for a quick adaptation based on some tens of training utterances. Two bounds are given for comparison with the results of the speaker adaptation, namely the recognition rates of speaker dependent and cross speaker recognition. Speaker dependent recognition is achieved by training the HMMs with nearly 1000 words uttered by the same speaker used in the tests. Cross speaker recognition, that gives a lower bound to the performance, concerns experiments in which the models were trained by a speaker different from that who uttered the test sentences. An adaptation algorithm using Parzen estimation and interpolation of the emission densities between the new and the old speaker models was investigated. It is able to give satisfactory recognition rates adapting the HMMs on the basis of only 40 training words uttered by the new speaker.
Key concepts: Speech recognition, Speaker recognition, Speaker diarisation, Hidden Markov model, Computer science, Adaptation (eye), Word error rate, Word (group theory)