1985•Unpublished venueRequires access

A theory and computational model of auditory monaural sound separation (stream, speech enhancement, selective attention, pitch perception, noise cancellation)

Mitchel Weintraub

Open publisher page 22 citations

Abstract

This thesis presents both a conceptual theory of how the auditory system uses monaural acoustic information to separate two simultaneous talkers and a computer model which is based on this theory. The information we believe the auditory system uses to separate sounds is reviewed, and a method to use this information for separating sounds is hypothesized. The computer model hypothesizes how the auditory system might use monaural acoustic information to determine how many sounds are present and the characteristics of each sound source. The use of periodicity information for the separation of sounds focuses on how experimental results about the neural encoding of sounds are combined with Licklider's theory of pitch perception. It is hypothesized that this acoustic information is interpreted by the auditory system using a two stage process: local regions in frequency and time are assigned to an intermediate representation called a group-object, and these group-objects are then subsequently assigned to a sound-stream. It is also hypothesized that the local encoding of periodicity information is used by the auditory system to assign local frequency-time regions to group-objects with similar features. A group-object is then assigned to one of the sound streams present based on the sound source that the auditory system believes generated this acoustic segment. A computer model which uses acoustic information to separate two simultaneous talkers is presented. The information present in the fine time stucture of a cochlear model's filterbank output is used as the input to the separation system. The computer model determines: how many people are speaking, whether each person's voice can be classified as periodic or nonperiodic, and what the spectral estimate for each talker is. The separation system uses both time and frequency continuity constraints in the modeling of each person's voice. Examples of how the system separates a male and female voice speaking a string of continuous digits are presented along with an evaluation of the current implementation.

About this research paper

What this paper is about

This thesis presents both a conceptual theory of how the auditory system uses monaural acoustic information to separate two simultaneous talkers and a computer model which is based on this theory. The information we believe the auditory system uses to separate sounds is reviewed, and a method to use this information for separating sounds is hypothesized. The computer model hypothesizes how the auditory system might use monaural acoustic information to determine how many sounds are present and the characteristics of each sound source. The use of periodicity information for the separation of sounds focuses on how experimental results about the neural encoding of sounds are combined with Licklider's theory of pitch perception. It is hypothesized that this acoustic information is interpreted by the auditory system using a two stage process: local regions in frequency and time are assigned to an intermediate representation called a group-object, and these group-objects are then subsequently assigned to a sound-stream. It is also hypothesized that the local encoding of periodicity information is used by the auditory system to assign local frequency-time regions to group-objects with similar features. A group-object is then assigned to one of the sound streams present based on the sound source that the auditory system believes generated this acoustic segment. A computer model which uses acoustic information to separate two simultaneous talkers is presented. The information present in the fine time stucture of a cochlear model's filterbank output is used as the input to the separation system. The computer model determines: how many people are speaking, whether each person's voice can be classified as periodic or nonperiodic, and what the spectral estimate for each talker is. The separation system uses both time and frequency continuity constraints in the modeling of each person's voice. Examples of how the system separates a male and female voice speaking a string of continuous digits are presented along with an evaluation of the current implementation.

Why it matters

OpenAlex reports 22 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

This thesis presents both a conceptual theory of how the auditory system uses monaural acoustic information to separate two simultaneous talkers and a computer model which is based on this theory. The information we believe the auditory system uses to separate sounds is reviewed, and a method to use this information for separating sounds is hypothesized. The computer model hypothesizes how the auditory system might use monaural acoustic information to determine how many sounds are present and the characteristics of each sound source. The use of periodicity information for the separation of sounds focuses on how experimental results about the neural encoding of sounds are combined with Licklider's theory of pitch perception. It is hypothesized that this acoustic information is interpreted by the auditory system using a two stage process: local regions in frequency and time are assigned to an intermediate representation called a group-object, and these group-objects are then subsequently assigned to a sound-stream. It is also hypothesized that the local encoding of periodicity information is used by the auditory system to assign local frequency-time regions to group-objects with similar features. A group-object is then assigned to one of the sound streams present based on the sound source that the auditory system believes generated this acoustic segment. A computer model which uses acoustic information to separate two simultaneous talkers is presented. The information present in the fine time stucture of a cochlear model's filterbank output is used as the input to the separation system. The computer model determines: how many people are speaking, whether each person's voice can be classified as periodic or nonperiodic, and what the spectral estimate for each talker is. The separation system uses both time and frequency continuity constraints in the modeling of each person's voice. Examples of how the system separates a male and female voice speaking a string of continuous digits are presented along with an evaluation of the current implementation.

Key concepts: Monaural, Computational auditory scene analysis, Auditory system, Auditory scene analysis, Computer science, Speech recognition, Acoustics, Precedence effect

Related papers

Back to paper searchBrowse research topicsOriginal source
A theory and computational model of auditory monaural sound separation (stream, speech enhancement, selective attention, pitch perception, noise cancellation) — Research Paper | ScholarLens