Scene-based Audio Implemented with Higher Order Ambisonics
Nils Peters, Deep Sen, Moo-Young Kim, Oliver Wuebbolt, S. Merrill Weiss
Abstract
Nils Peters, Deep Sen, Moo-Young Kim, Oliver Wuebbolt, S. Merrill Weiss
Abstract
Scene-based audio uses a sound-field technology called “higher-order ambisonics” (HOA) to create holistic descriptions of both live-captured and artistically created sound scenes that are independent of specific loudspeaker layouts. For efficient representation, the audio can be carried as a set of PCM channels that contain predominant sounds and ambience in separate tracks. Standard audio bandwidth-compression techniques can then be applied to the PCM channels. This approach is in contrast to conventional channel-based sound representations in which one signal is used for each loudspeaker of a target reproduction system, with the implication that upmixing or downmixing is required when loudspeaker configurations other than the intended one are used for actual reproduction. This paper examines how, with scene-based audio, there can be satisfactory reproduction of immersive sound at bitrates corresponding to the equivalent of only six channels, whereas an alternative sound-field method that exclusively employs audio objects typically involves much higher bitrates.
OpenAlex reports 8 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Scene-based audio uses a sound-field technology called “higher-order ambisonics” (HOA) to create holistic descriptions of both live-captured and artistically created sound scenes that are independent of specific loudspeaker layouts. For efficient representation, the audio can be carried as a set of PCM channels that contain predominant sounds and ambience in separate tracks. Standard audio bandwidth-compression techniques can then be applied to the PCM channels. This approach is in contrast to conventional channel-based sound representations in which one signal is used for each loudspeaker of a target reproduction system, with the implication that upmixing or downmixing is required when loudspeaker configurations other than the intended one are used for actual reproduction. This paper examines how, with scene-based audio, there can be satisfactory reproduction of immersive sound at bitrates corresponding to the equivalent of only six channels, whereas an alternative sound-field method that exclusively employs audio objects typically involves much higher bitrates.
Key concepts: Ambisonics, Loudspeaker, Surround sound, Sound recording and reproduction, Computer science, Audio signal, Digital audio, Sound quality