The objective PSNR metric is known to correlate quite poorly with subjective assessments of video coding quality. Thus, a number of alternative VQA measures such as (MS-)SSIM and VMAF have been proposed. These, however, are often algorithmically complex and difficult to use for visually motivated encoder optimization tasks, especially subjectively optimized bit allocation. In this paper we show that, by way of low-complexity enhancements of our previous work on a perceptually weighted PSNR (WPSNR) metric, addressing shortcomings with video and ultra high-definition content, the prediction of human judgments of video coding quality by the WPSNR can be improved. In fact, the resulting XPSNR seems to match the performance of the aforementioned state-of-the-art methods.
Video coding in the YCbCr color space has been widely used, since it is efficient for compression, but it can result in color distortion due to conversion error. Meanwhile, coding in the RGB color space maintains high color fidelity, having the drawback of a substantial bitrate increase with respect to YCbCr coding. Cross-component prediction (CCP) efficiently compresses video content by decorrelating color components while keeping high color fidelity. In this scheme, the chroma residual signal is predicted from the luma residual signal inside the coding loop. This paper gives a description of the CCP scheme from several points of view, from theoretical background to practical implementation. The proposed CCP scheme has been evaluated in standardization communities and adopted into H.265/High Efficiency Video Coding (HEVC) Range Extensions. The experimental results show significant coding performance improvements for both natural and screen content video, while the quality of all color components is maintained. The average coding gains for natural video are 17% and 5% bitrate reduction in the case of intra coding and 11% and 4% in the case of inter coding for RGB and YCbCr coding, respectively, while the average increment of encoding and decoding times in the HEVC reference software implementation are 10% and 4%, respectively.
In this paper, we describe a video coding design that enables a higher coding efficiency than the HEVC standard. The proposed video codec follows the design of block-based hybrid video coding, but includes a number of advanced coding tools. A part of the incorporated advanced concepts was developed by the Joint Video Exploration Team, while others are newly proposed. The key aspects of these newly proposed tools are the following. A video frame is subdivided into rectangles of variable size using a binary partitioning with variable split ratios. Three new approaches for generating spatial intra prediction signals are supported: A line-wise application of conventional intra prediction modes, coupled with a mode-dependent processing order, a region-based template matching prediction method and intra prediction modes based on neural networks. For motion-compensated prediction, a multi-hypothesis mode with more than two motion hypotheses can be used. In transform coding, mode dependent combinations of primary and secondary transforms are applied. Moreover, scalar quantization is replaced by trellis-coded quantization and the entropy coding of the quantized transform coefficients is improved. The intra and inter prediction signals can be filtered using an edge-preserving diffusion filter or a non-linear DCT-based thresholding operation. The video codec includes an adaptive in-loop filter for which one of three classifiers can be chosen on a picture basis. We also incorporated an optional encoder control, which adjusts the quantization parameters based on a perceptually motivated distortion measure. In a random access scenario, our proposed video codec achieves luma BD-rate savings between 32.5% for HDR HLG UHD and 39.6% for SDR UHD over the HEVC (HM software) anchor for different categories of test sequences.
In state-of-the-art video compression residual coding is done by transforming the prediction error signals into a less correlated representation and performing the quantization and entropy coding in the transform domain. For complexity reasons usually separable transforms are used. A more flexible transform structure is given by row-column transforms, which apply a separate transform to each row and each column of a signal block. This paper describes a method for training such structured transforms by maximizing the data likelihood under a parameterized probabilistic model with a compelled structure. An explicit model is derived for the case of row-column transforms and its efficiency is demonstrated in the application of video compression. It is shown that trained row-column transforms achieve almost the same coding gain as unconstrained KLTs when applied as secondary transforms, while the encoder and decoder runtime are the same as in the separable transform case.
It is well known that input-invariant quantization in perceptual image or video coding often leads to visually suboptimal results and that quantization parameter adaptation (QPA) based on a model of the human visual system can improve subjective coding quality. This paper introduces a simple low-complexity QPA algorithm, controlled using a block-wise perceptually weighted distortion measure representing a generalization of the PSNR metric. The weighting scheme of this WPSNR metric is based on a psychovisual model. It directly leads to a perceptually adapted scaling of the block-wise Lagrange parameter used in the bit-allocation process in the encoder and, consequently, to a block-wise QPA. Unlike prior QPA approaches, the proposal avoids classifications of picture regions and easily extends from still-image or grayscale to video or chromatic coding. The WPSNR metric also uses fewer algorithmic operations than e. g. the multiscale structural similarity measure (MS-SSIM). Due to the results of two formal subjective tests indicating its visual benefit, the QPA proposal has been adopted into VTM, the currently developed Versatile Video Coding (VVC) reference software.
Bit-allocation based on the MSE is computationally convenient in image and video compression, but leads to perceptually suboptimal compression results. Distortion sensitivity, modeled as a reference specific property, can be used to improve the accuracy of perceptual quality prediction based on the MSE. This paper shows how distortion sensitivity directly leads to computationally beneficial perceptual optimization of irrelevance reduction and, thereby, of bit-allocation in image and video compression. To this end distortion sensitivity is estimated using a deep convolutional neural network. The proposed method of distortion sensitive bit-allocation is evaluated experimentally using HEVC and on our testset shows average bit-rate reductions with regard to the MOS of 15.9% compared to constant QP-based bit-allocation and 7.3% compared to state-of-the-art perceptual bit-allocation schemes.
In this paper we combine state of the art video compression and Partial Differential Equation (PDE) based image processing methods. We introduce a new signal adaptive method to filter the predictions of a hybrid video codec using a system of PDEs describing a diffusion process. The method can be applied to intra as well as inter predictions. The filter is embedded into the framework of HEVC. The efficiency of the HEVC video codec is improved by up to -2.76% for All Intra and -3.56% for Random Access measured in Bjøntegaard delta (BD) rate. Coding gains of up to -8.76% can be observed for individual test sequences.
The support for chroma formats other than 4:2:0 and bit depths higher than 10 bit has been included into the High Efficiency Video Coding (HEVC) standard by the publication of Version 2, In contrast to 4:2:2 profiles without additional dedicated coding tools relative to the Main profile, all 4:4:4 profiles specify a novel compression efficiency tool referred to as Cross-Component Prediction (CCP). This paper describes and analyses two additional extensions relative to the specified CCP variant in Version 2. In the first extension, an additional predictor is introduced, particularly, the first chroma component can also serve as the predictor for the second chroma component beside the luma component. The second extension is concerned with the binarization of the single prediction model parameter, which usually has a different distribution for R'G'B' than for Y'CbCr content. Experimental results show that the in this paper presented extensions can improve the compression efficiency by -0.7% on average for R'G'B' content, and by -1.1% for Y'CbCr content, respectively, both in terms of BD rate.
Although being one of the simplest and most widely used image quality metrics (IQMs) the peak signal-to-noise ratio (PSNR) correlates only poorly with visual quality as perceived by humans. Based on an analysis of the non-linear mapping from PSNR to mean opinion scores (MOS) we identify a functional mapping parameter to adapt the PSNR perceptually meaningful. Neurophysiologically motivated, a shearlet-based correction is proposed for controlling this perceptual PSNR adaption. The performance of the proposed perceptually adapted PSNR is evaluated on the LIVE and TID2013 databases and shows to be superior or comparable to benchmark IQMs.
This paper proposes a reduced reference image quality assessment method using only a low number of features. It involves a shearlet decomposition, directional pooling of the obtained coefficient and extracts the scalewise statistical location parameter as a feature. The proposed method is tested and compared to similar approaches on the LIVE image database. On this database it outperforms the compared methods on five of seven distortion types and on the full testset with a linear correlation of = 0.89.
Three experiments addressing the assessment of perceived image quality in a patch-based manner are compared for HEVC compression artifacts. It is shown that image patches of a size small as 128×128 pixel are large enough to evaluate the perceived image quality in a Degradation Category Rating (DCR) setting. Ratings obtained with 128×128 pixel sized images patches and 512×512 pixel sized images of the same spatial statistics show a correlation of r=0.99. Based on this finding, image quality assessment of 128×128 pixel sized image patches degraded by HEVC compression is compared for controlled lab environment and uncontrolled crowdsourcing settings. Although we find high overall correlation between the quality ratings obtained in the two environments, observers tend to give worse ratings in the crowdsourcing setting and for conditions of higher quality a reduction of correlation is observed. These findings have implications for choosing controlled vs. uncontrolled viewing conditions for image quality assessment for real-life applications.
The Marie Skłodowska-Curie Initial Training Network on Perceptually Optimised Video Compression, called PROVISION, is a collaborative network between academic and industrial partners within the 7th Framework Programme of the European Commission. Its key aim is to deliver technical advances well beyond the capabilities of current video compression standards, focusing on human visual perception for the future of video compression. This paper gives an overview of the PROVISION project, including the challenges that motivated its creation, the related scientific work supporting the project and its main objectives. Examples of some specific research topics addressed in the context of the project are also presented.
Due to the higher requirements associated with Ultra High Definition (UHD) resolutions in terms of memory and transmission bandwidth, the feasibility of UHD video communication applications is strongly dependent on the performance of video compression solutions. Even though the High Efficiency Video Coding (HEVC) standard allows significantly superior rate-distortion performances compared to previous video coding standards, further performance improvements are possible when exploiting the perceptual properties of the Human Visual System (HVS). This paper proposes a novel perceptual-based solution fully compliant with the HEVC standard, where a low complexity Just Noticeable Distortion model is used to drive the encoder's rate-distortion optimised quantisation process. This technique allows a simple and effective way to influence the decisions made at the encoder, based on the limitations of the HVS. The experiments conducted for UHD resolutions show average bitrate savings of 21% with no visual quality degradations when compared to the HEVC reference software.
The work on support for higher bit depths and 4:2:2 as well as 4:4:4 chroma sampling formats for the High Efficiency Video Coding (HEVC) standard is currently being conducted under the term Range Extensions (RExt). A technique that exploits the correlation between residual color components in 4:4:4 chroma sampling format, also referred to as Cross-Component Prediction (CCP), has been adopted as part of the current RExt draft. In this paper, this relatively simple but yet effective CCP scheme is presented. Conceptually, CCP relies on the idea that an adaptively switched predictor based on a linear model is invoked for coding of the residuals of the second and third color component by using the residual of the first color component. Experimental results show that, depending on the underlying color space, average bit-rate savings in the range of 2-18% or 3-26% can be achieved by CCP for test sets of natural and screen content, respectively.