2008•Unpublished venueOpen access

Scalable 3-D wavelet video coding

Wenbo Zong

Open full text 0 citations

Abstract

Concerning the various applications that involve heterogeneous networks, computing capabilities, and/or display terminals, such as Internet video streaming and live video surveillance, a scalable video coding system provides an effective solution to address this issue by offering some scalable functionalities including temporal scalability, spatial scalability and signal-to-noise ratio (SNR) scalability, while achieving high compression performance.Conventional hybrid coders like MPEG-2 and MPEG-4 are not able to offer scalability efficiently due to the recursive motion-compensated (MC) prediction loop.In recent years, motion-compensated wavelet coding (MCWC) has emerged as a promising video coding approach to offering both high compression performance and versatile scalabilities.In this research, we first introduce the 3-D MCWC technology and its three possible structures: t+2D, 2D+t, and 2D+t+2D.We then investigate two issues related to MCWC.The first issue is the spatial scalability of the t+2D scheme, which suffers from degraded rate-distortion (R-D) performance for lower spatial resolutions.We investigate the problem from both the decoder side and the encoder side.At the decoder side, we present an in-depth analysis on the inverse motioncompensated temporal filtering (MCTF), or IMCTF, when reconstructing lower spatial resolution videos, and propose a method to improve the IMCTF process at the decoder.At the encoder side, we analyze the MCTF process in detail, and highlight the root source of aliasing that is not cancellable for lower spatial resolutions at the decoder.We propose a scheme named low-to-high lifting MCTF (LTH-MCTF) to eliminate this aliasing.Furthermore, we propose a practical solution, named pyramidal t+2D, that applies t+2D coding to each level of the frame pyramids to produce one separate bitstream with optimal R-D performance for each spatial resolution.The second issue we investigate in this research is the possibility of applying set-partitioning techniques to the temporal high-pass frames generated by variable size block motion compensation.We observe that the MC prediction residues exhibit different statistics in motion blocks of different sizes.We thus propose to partition the wavelet coefficients into groups according to the size of the motion block to which they correspond, which is shown to reduce the entropy with bitplane quantization.also be decoded at a lower quality, resolution and/or frame rate according to the requirements (e.g., channel capacity, storage capacity, processing power, display size).This kind of flexibility of partially decoding an encoded bitstream to meet the heterogeneous requirements of various applications is achieved in a limited way in conventional hybrid coders, such as MPEG-2 and MPEG-4.In the following section, we shall briefly review the issues of scalability encountered in the conventional hybrid coders, especially on their SNR (quality) scalability.This would clearly establish concrete justification for the necessity of investigating and designing a scalable coder.

Open-access reader

About this research paper

What this paper is about

Concerning the various applications that involve heterogeneous networks, computing capabilities, and/or display terminals, such as Internet video streaming and live video surveillance, a scalable video coding system provides an effective solution to address this issue by offering some scalable functionalities including temporal scalability, spatial scalability and signal-to-noise ratio (SNR) scalability, while achieving high compression performance.Conventional hybrid coders like MPEG-2 and MPEG-4 are not able to offer scalability efficiently due to the recursive motion-compensated (MC) prediction loop.In recent years, motion-compensated wavelet coding (MCWC) has emerged as a promising video coding approach to offering both high compression performance and versatile scalabilities.In this research, we first introduce the 3-D MCWC technology and its three possible structures: t+2D, 2D+t, and 2D+t+2D.We then investigate two issues related to MCWC.The first issue is the spatial scalability of the t+2D scheme, which suffers from degraded rate-distortion (R-D) performance for lower spatial resolutions.We investigate the problem from both the decoder side and the encoder side.At the decoder side, we present an in-depth analysis on the inverse motioncompensated temporal filtering (MCTF), or IMCTF, when reconstructing lower spatial resolution videos, and propose a method to improve the IMCTF process at the decoder.At the encoder side, we analyze the MCTF process in detail, and highlight the root source of aliasing that is not cancellable for lower spatial resolutions at the decoder.We propose a scheme named low-to-high lifting MCTF (LTH-MCTF) to eliminate this aliasing.Furthermore, we propose a practical solution, named pyramidal t+2D, that applies t+2D coding to each level of the frame pyramids to produce one separate bitstream with optimal R-D performance for each spatial resolution.The second issue we investigate in this research is the possibility of applying set-partitioning techniques to the temporal high-pass frames generated by variable size block motion compensation.We observe that the MC prediction residues exhibit different statistics in motion blocks of different sizes.We thus propose to partition the wavelet coefficients into groups according to the size of the motion block to which they correspond, which is shown to reduce the entropy with bitplane quantization.also be decoded at a lower quality, resolution and/or frame rate according to the requirements (e.g., channel capacity, storage capacity, processing power, display size).This kind of flexibility of partially decoding an encoded bitstream to meet the heterogeneous requirements of various applications is achieved in a limited way in conventional hybrid coders, such as MPEG-2 and MPEG-4.In the following section, we shall briefly review the issues of scalability encountered in the conventional hybrid coders, especially on their SNR (quality) scalability.This would clearly establish concrete justification for the necessity of investigating and designing a scalable coder.

Why it matters

A significance statement is not available in the OpenAlex record.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Concerning the various applications that involve heterogeneous networks, computing capabilities, and/or display terminals, such as Internet video streaming and live video surveillance, a scalable video coding system provides an effective solution to address this issue by offering some scalable functionalities including temporal scalability, spatial scalability and signal-to-noise ratio (SNR) scalability, while achieving high compression performance.Conventional hybrid coders like MPEG-2 and MPEG-4 are not able to offer scalability efficiently due to the recursive motion-compensated (MC) prediction loop.In recent years, motion-compensated wavelet coding (MCWC) has emerged as a promising video coding approach to offering both high compression performance and versatile scalabilities.In this research, we first introduce the 3-D MCWC technology and its three possible structures: t+2D, 2D+t, and 2D+t+2D.We then investigate two issues related to MCWC.The first issue is the spatial scalability of the t+2D scheme, which suffers from degraded rate-distortion (R-D) performance for lower spatial resolutions.We investigate the problem from both the decoder side and the encoder side.At the decoder side, we present an in-depth analysis on the inverse motioncompensated temporal filtering (MCTF), or IMCTF, when reconstructing lower spatial resolution videos, and propose a method to improve the IMCTF process at the decoder.At the encoder side, we analyze the MCTF process in detail, and highlight the root source of aliasing that is not cancellable for lower spatial resolutions at the decoder.We propose a scheme named low-to-high lifting MCTF (LTH-MCTF) to eliminate this aliasing.Furthermore, we propose a practical solution, named pyramidal t+2D, that applies t+2D coding to each level of the frame pyramids to produce one separate bitstream with optimal R-D performance for each spatial resolution.The second issue we investigate in this research is the possibility of applying set-partitioning techniques to the temporal high-pass frames generated by variable size block motion compensation.We observe that the MC prediction residues exhibit different statistics in motion blocks of different sizes.We thus propose to partition the wavelet coefficients into groups according to the size of the motion block to which they correspond, which is shown to reduce the entropy with bitplane quantization.also be decoded at a lower quality, resolution and/or frame rate according to the requirements (e.g., channel capacity, storage capacity, processing power, display size).This kind of flexibility of partially decoding an encoded bitstream to meet the heterogeneous requirements of various applications is achieved in a limited way in conventional hybrid coders, such as MPEG-2 and MPEG-4.In the following section, we shall briefly review the issues of scalability encountered in the conventional hybrid coders, especially on their SNR (quality) scalability.This would clearly establish concrete justification for the necessity of investigating and designing a scalable coder.

Key concepts: Computer science, Wavelet, Coding (social sciences), Scalability, Speech recognition, Cognitive science, Data science, Psychology

Related papers

Back to paper searchBrowse research topicsOriginal source
Scalable 3-D wavelet video coding — Research Paper | ScholarLens