A warped linear-prediction-based subband audio coding algorithm
Yu Rongshan, C.C. Ko
Abstract
Yu Rongshan, C.C. Ko
Abstract
A novel audio coding algorithm is proposed where the warped-linear prediction (WLP) technique is employed to construct a perceptual pre- and post-filter for subband audio coding. A modified signal-to-mask ratio (SMR) calculation is given for subband coding of the WLP residuals of audio signals. The concept of perceptual entropy (PE) is extended to subband coding, resulting in the subband perceptual entropy (SPE), which gives a short time estimate of the lowest possible bit rate for transparent subband audio coding. Two WLP models with frequency responses approximating the spectral shape of the masking threshold are investigated and it is found that the residual signals of both models contain less SPE compared with that of the original audio signals. Subjective tests show that the proposed audio codec operating at 56 kbps has a perceptual quality comparable to MPEG-1 audio Layer II operating at 64 kbps.
OpenAlex reports 11 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
A novel audio coding algorithm is proposed where the warped-linear prediction (WLP) technique is employed to construct a perceptual pre- and post-filter for subband audio coding. A modified signal-to-mask ratio (SMR) calculation is given for subband coding of the WLP residuals of audio signals. The concept of perceptual entropy (PE) is extended to subband coding, resulting in the subband perceptual entropy (SPE), which gives a short time estimate of the lowest possible bit rate for transparent subband audio coding. Two WLP models with frequency responses approximating the spectral shape of the masking threshold are investigated and it is found that the residual signals of both models contain less SPE compared with that of the original audio signals. Subjective tests show that the proposed audio codec operating at 56 kbps has a perceptual quality comparable to MPEG-1 audio Layer II operating at 64 kbps.
Key concepts: Sub-band coding, Speech coding, Speech recognition, Codec, Computer science, Linear prediction, Audio signal, Adaptive Multi-Rate audio codec