2017Unpublished venueRequires access

Simple and Effective Speech Enhancement for Visual Microphone

Juhyun Ahn, Dai‐Jin Kim

Open publisher page 6 citations

Abstract

Visual microphone is a technique that recovers the sound from a silent video. The simplest way to improve sound recovery performance of the visual microphone is by applying the traditional speech enhancement algorithms which are based on complicated filter designs or sound models. This paper proposes a simple and effective speech enhancement for visual microphone (SEVM) that suppress spectrum components with small amplitude than a predefined threshold value, which exploits the unique properties that the sound spectrum recovered from the visual microphone is relatively high and the noise spectrum generated motion estimation error and damped oscillation is relatively low. The proposed SEVM method can also be easily extended to a multichannel case that multiple speech signals are recovered from multiple cameras. Experimental results show the proposed SEVM method better performance than the traditional speech enhancement algorithms in terms of log-likelihood ratio (LLR), signal to noise ratio (SNR), segmental SNR (SegSNR) and cepstral distance (CEP). From these results, we convince that the proposed SEVM method that is adapted to the visual microphone is really simple and effective than the traditional speech enhancement methods that are just extended to the visual microphone as a post-processing.

About this research paper

What this paper is about

Visual microphone is a technique that recovers the sound from a silent video. The simplest way to improve sound recovery performance of the visual microphone is by applying the traditional speech enhancement algorithms which are based on complicated filter designs or sound models. This paper proposes a simple and effective speech enhancement for visual microphone (SEVM) that suppress spectrum components with small amplitude than a predefined threshold value, which exploits the unique properties that the sound spectrum recovered from the visual microphone is relatively high and the noise spectrum generated motion estimation error and damped oscillation is relatively low. The proposed SEVM method can also be easily extended to a multichannel case that multiple speech signals are recovered from multiple cameras. Experimental results show the proposed SEVM method better performance than the traditional speech enhancement algorithms in terms of log-likelihood ratio (LLR), signal to noise ratio (SNR), segmental SNR (SegSNR) and cepstral distance (CEP). From these results, we convince that the proposed SEVM method that is adapted to the visual microphone is really simple and effective than the traditional speech enhancement methods that are just extended to the visual microphone as a post-processing.

Why it matters

OpenAlex reports 6 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Visual microphone is a technique that recovers the sound from a silent video. The simplest way to improve sound recovery performance of the visual microphone is by applying the traditional speech enhancement algorithms which are based on complicated filter designs or sound models. This paper proposes a simple and effective speech enhancement for visual microphone (SEVM) that suppress spectrum components with small amplitude than a predefined threshold value, which exploits the unique properties that the sound spectrum recovered from the visual microphone is relatively high and the noise spectrum generated motion estimation error and damped oscillation is relatively low. The proposed SEVM method can also be easily extended to a multichannel case that multiple speech signals are recovered from multiple cameras. Experimental results show the proposed SEVM method better performance than the traditional speech enhancement algorithms in terms of log-likelihood ratio (LLR), signal to noise ratio (SNR), segmental SNR (SegSNR) and cepstral distance (CEP). From these results, we convince that the proposed SEVM method that is adapted to the visual microphone is really simple and effective than the traditional speech enhancement methods that are just extended to the visual microphone as a post-processing.

Key concepts: Microphone, Speech enhancement, Computer science, Speech recognition, Noise-canceling microphone, Noise (video), Cepstrum, Filter (signal processing)

Related papers

Back to paper searchBrowse research topicsOriginal source
Simple and Effective Speech Enhancement for Visual Microphone — Research Paper | ScholarLens