2008•Unpublished venueRequires access

Humanoid separation of speech sources in reverberant environments

Sylvia Schulz, Thorsten Herfet

Open publisher page 2 citations

Abstract

This paper presents a framework for humanoid binaural source separation in a non-ideal reverberant environment and investigates the ideal position of a human head for several separation algorithms. A movable human dummy head residing in a normal office room is used to model the conditions humans experience in a complex auditory scene. Prior to source separation the dummy head analyzes the auditory scene, estimates several parameters and aligns to the best separation position. The estimated information is used to enhance the succeeding source separation: Several parallel separation paths estimate independent time-frequency masks to demix the target source. A combination stage infers a final estimate of the ideal binary mask. The presented approach outperforms fixed beamforming and the DUET source separation algorithm in reverberant auditory scenes consisting of two and three speech sources by up to 17 dB SIR gain.

About this research paper

What this paper is about

This paper presents a framework for humanoid binaural source separation in a non-ideal reverberant environment and investigates the ideal position of a human head for several separation algorithms. A movable human dummy head residing in a normal office room is used to model the conditions humans experience in a complex auditory scene. Prior to source separation the dummy head analyzes the auditory scene, estimates several parameters and aligns to the best separation position. The estimated information is used to enhance the succeeding source separation: Several parallel separation paths estimate independent time-frequency masks to demix the target source. A combination stage infers a final estimate of the ideal binary mask. The presented approach outperforms fixed beamforming and the DUET source separation algorithm in reverberant auditory scenes consisting of two and three speech sources by up to 17 dB SIR gain.

Why it matters

OpenAlex reports 2 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

This paper presents a framework for humanoid binaural source separation in a non-ideal reverberant environment and investigates the ideal position of a human head for several separation algorithms. A movable human dummy head residing in a normal office room is used to model the conditions humans experience in a complex auditory scene. Prior to source separation the dummy head analyzes the auditory scene, estimates several parameters and aligns to the best separation position. The estimated information is used to enhance the succeeding source separation: Several parallel separation paths estimate independent time-frequency masks to demix the target source. A combination stage infers a final estimate of the ideal binary mask. The presented approach outperforms fixed beamforming and the DUET source separation algorithm in reverberant auditory scenes consisting of two and three speech sources by up to 17 dB SIR gain.

Key concepts: Binaural recording, Source separation, Separation (statistics), Computer science, Position (finance), Ideal (ethics), Binary number, Speech recognition

Related papers

Back to paper searchBrowse research topicsOriginal source
Humanoid separation of speech sources in reverberant environments — Research Paper | ScholarLens