2011•Intelligent Computer and ApplicationsRequires access

The Voice Activity Detection Algorithm based on the Mean of Long and Short Time Frame Energy

Shiwen Deng

Open publisher page 0 citations

Abstract

Aimed to remove the effect of the non-stationary noise in speech processing system,the paper proposes a voice activity detection(VAD) algorithm which meanly uses the mean of long and short time frame energy.The proposed VAD algorithm is based on two hypothesis,the first one is that,when the signal is decomposed by sparse representation plus the set of speech underlying structures,the part of energy of speech will be effectively reserved,while the part of energy of non-speech will be partly removed;the other one is that,when time-domain signal is reconstructed from the sparse coefficients which are generated by the above-mentioned sparse representation,the energy of the part of speech is larger than the part of non-speech among the reconstructed signal.In the proposed VAD algorithm,the energy of the reconstructed signal is used as the audio feature,the Simultaneously,taking the current frame as the center,mean of a certain number of adjacent shout-time frame energy is used as the score of the current frame;the mean of fixed number of long-time frame energy before the current frame is considered as the detection threshold,and then the current frame is claimed as speech when the score of current frame is larger than the detection threshold,or else it is claimed as non-speech.The experimental results show that the proposed VAD algorithm can distinguish the speech and non-speech effectively,and its performance outperforms the Gaussian distribution based likelihood ratio test(LRT_Gaussian) algorithm.

About this research paper

What this paper is about

Aimed to remove the effect of the non-stationary noise in speech processing system,the paper proposes a voice activity detection(VAD) algorithm which meanly uses the mean of long and short time frame energy.The proposed VAD algorithm is based on two hypothesis,the first one is that,when the signal is decomposed by sparse representation plus the set of speech underlying structures,the part of energy of speech will be effectively reserved,while the part of energy of non-speech will be partly removed;the other one is that,when time-domain signal is reconstructed from the sparse coefficients which are generated by the above-mentioned sparse representation,the energy of the part of speech is larger than the part of non-speech among the reconstructed signal.In the proposed VAD algorithm,the energy of the reconstructed signal is used as the audio feature,the Simultaneously,taking the current frame as the center,mean of a certain number of adjacent shout-time frame energy is used as the score of the current frame;the mean of fixed number of long-time frame energy before the current frame is considered as the detection threshold,and then the current frame is claimed as speech when the score of current frame is larger than the detection threshold,or else it is claimed as non-speech.The experimental results show that the proposed VAD algorithm can distinguish the speech and non-speech effectively,and its performance outperforms the Gaussian distribution based likelihood ratio test(LRT_Gaussian) algorithm.

Why it matters

A significance statement is not available in the OpenAlex record.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Aimed to remove the effect of the non-stationary noise in speech processing system,the paper proposes a voice activity detection(VAD) algorithm which meanly uses the mean of long and short time frame energy.The proposed VAD algorithm is based on two hypothesis,the first one is that,when the signal is decomposed by sparse representation plus the set of speech underlying structures,the part of energy of speech will be effectively reserved,while the part of energy of non-speech will be partly removed;the other one is that,when time-domain signal is reconstructed from the sparse coefficients which are generated by the above-mentioned sparse representation,the energy of the part of speech is larger than the part of non-speech among the reconstructed signal.In the proposed VAD algorithm,the energy of the reconstructed signal is used as the audio feature,the Simultaneously,taking the current frame as the center,mean of a certain number of adjacent shout-time frame energy is used as the score of the current frame;the mean of fixed number of long-time frame energy before the current frame is considered as the detection threshold,and then the current frame is claimed as speech when the score of current frame is larger than the detection threshold,or else it is claimed as non-speech.The experimental results show that the proposed VAD algorithm can distinguish the speech and non-speech effectively,and its performance outperforms the Gaussian distribution based likelihood ratio test(LRT_Gaussian) algorithm.

Key concepts: Computer science, Frame (networking), Energy (signal processing), Speech recognition, Algorithm, SIGNAL (programming language), Feature (linguistics), Voice activity detection

Related papers

Back to paper searchBrowse research topicsOriginal source
The Voice Activity Detection Algorithm based on the Mean of Long and Short Time Frame Energy — Research Paper | ScholarLens