Scene-Based Audio Implemented with Higher Order Ambisonics (HOA)
Nils Peters, Deep Sen, Moo-Young Kim, Oliver Wuebbolt, S. Merrill Weiss
Abstract
Nils Peters, Deep Sen, Moo-Young Kim, Oliver Wuebbolt, S. Merrill Weiss
Abstract
Scene-based Audio uses a sound-field technology called “Higher Order Ambisonics” (HOA) to create holistic descriptions of both live-captured and artistically-created sound scenes that are independent of specific loudspeaker layouts. For efficient representation, the audio can be carried as a set of PCM channels that contain predominant sounds and ambience in separate tracks. Standard audio bandwidth-compression techniques then can be applied to the PCM channels. This approach is in contrast to conventional channel-based sound representations, in which one signal is used for each loudspeaker of a target reproduction system, with the implication that upmixing or downmixing is required when loudspeaker configurations other than the intended one are used for actual reproduction. This paper examines how, with Scene-based Audio, there can be satisfactory reproduction of immersive sound at bitrates corresponding to the equivalent of just 6 channels, while an alternative sound-field method that exclusively employs audio objects typically involves much higher bitrates.
OpenAlex reports 5 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Scene-based Audio uses a sound-field technology called “Higher Order Ambisonics” (HOA) to create holistic descriptions of both live-captured and artistically-created sound scenes that are independent of specific loudspeaker layouts. For efficient representation, the audio can be carried as a set of PCM channels that contain predominant sounds and ambience in separate tracks. Standard audio bandwidth-compression techniques then can be applied to the PCM channels. This approach is in contrast to conventional channel-based sound representations, in which one signal is used for each loudspeaker of a target reproduction system, with the implication that upmixing or downmixing is required when loudspeaker configurations other than the intended one are used for actual reproduction. This paper examines how, with Scene-based Audio, there can be satisfactory reproduction of immersive sound at bitrates corresponding to the equivalent of just 6 channels, while an alternative sound-field method that exclusively employs audio objects typically involves much higher bitrates.
Key concepts: Ambisonics, Computer science, Order (exchange), Computer graphics (images), Loudspeaker, Engineering, Finance, Economics