Speaker Verification Based on Different Vector Quantization Techniques with Gaussian Mixture Models
Sheeraz Memon, Margaret Lech, Namunu C. Maddage
Abstract
Sheeraz Memon, Margaret Lech, Namunu C. Maddage
Abstract
The introduction of Gaussian mixture models (GMMs) in the field of speaker verification has led to very good results. This paper illustrates an evolution in state-of-the-art speaker verification by highlighting the contribution of recently established information theoretic based vector quantization technique. We explore the novel application of three different vector quantization algorithms, namely K-means, Linde-Buzo-Gray (LBG) and information theoretic vector quantization (ITVQ) for efficient speaker verification. The expectation maximization (EM) algorithm used by GMM requires a prohibitive amount of iterations to converge. In this paper, comparable alternatives to EM including K-means, LBG and ITVQ algorithm were tested. The GMM-ITVQ algorithm was found to be the most efficient alternative for the GMM-EM. It gives correct classification rates at a similar level to that of GMM-EM. Finally, representative performance benchmarks and system behaviour experiments on NIST SRE corpora are presented.
OpenAlex reports 20 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
The introduction of Gaussian mixture models (GMMs) in the field of speaker verification has led to very good results. This paper illustrates an evolution in state-of-the-art speaker verification by highlighting the contribution of recently established information theoretic based vector quantization technique. We explore the novel application of three different vector quantization algorithms, namely K-means, Linde-Buzo-Gray (LBG) and information theoretic vector quantization (ITVQ) for efficient speaker verification. The expectation maximization (EM) algorithm used by GMM requires a prohibitive amount of iterations to converge. In this paper, comparable alternatives to EM including K-means, LBG and ITVQ algorithm were tested. The GMM-ITVQ algorithm was found to be the most efficient alternative for the GMM-EM. It gives correct classification rates at a similar level to that of GMM-EM. Finally, representative performance benchmarks and system behaviour experiments on NIST SRE corpora are presented.
Key concepts: Vector quantization, Mixture model, Speaker verification, Computer science, Quantization (signal processing), NIST, Gaussian, Expectation–maximization algorithm