The complexity of modern codecs along with the increased need of delivering high-quality videos at low bitrates has reinforced the idea of a per-clip tailoring of parameters for optimised rate-distortion performance. While the objective quality metrics used for Standard Dynamic Range (SDR) videos have been well studied, the transitioning of consumer displays to support High Dynamic Range (HDR) videos, poses a new challenge to rate-distortion optimisation. In this paper, we review the popular HDR metrics DeltaE100 (DE100), PSNRL100, wPSNR, and HDR-VQM. We measure the impact of employing these metrics in per-clip direct search optimisation of the rate-distortion Lagrange multiplier in AV1. We report, on 35 HDR videos, average Bjontegaard Delta Rate (BD-Rate) gains of 4.675%, 2.226%, and 7.253% in terms of DE100, PSNRL100, and HDR-VQM. We also show that the inclusion of chroma in the quality metrics has a significant impact on optimisation, which can only be partially addressed by the use of chroma offsets.
Video compression is complicated by degradation in User Generated Content (UGC). Preprocessing the data before encoding improves compression. However the impact of the preprocessor depends not only on the codec and the filter strength of the preprocessor being used but also on the target bitrate of the encode and the level of degradation. In this paper we present a framework for modelling this relationship and estimating the optimal filter strength for a particular codec/preprocessor/bitrate/degradation combination. We examine two preprocessors based on classical and DNN ideas, and two codecs AV1, VP9. We find that up to 2dB of quality gain can result from preprocessing at constant bitrate and our estimator is accurate enough to capture most of these gains.
Since the adoption of VP9 by Netflix in 2016, royalty-free coding standards continued to gain prominence through the activities of the AOMedia consortium. AV1, the latest open source standard, is now widely supported. In the early years after standardisation, HDR video tends to be under served in open source encoders for a variety of reasons including the relatively small amount of true HDR content being broadcast and the challenges in RD optimisation with that material. AV1 codec optimisation has been ongoing since 2020 including consideration of the computational load. In this paper, we explore the idea of direct optimisation of the Lagrangian λ parameter used in the rate control of the encoders to estimate the optimal Rate-Distortion trade-off achievable for a High Dynamic Range signalled video clip. We show that by adjusting the Lagrange multiplier in the RD optimisation process on a frame-hierarchy basis, we are able to increase the Bjontegaard difference rate gains by more than 3.98× on average without visually affecting the quality.
Recent advances have shown that latent representations of pre-trained Deep Convolutional Neural Networks (DCNNs) for classification can be modelled to generate scores that are well correlated with human perceptual judgement. In this paper we seek to extend the use of perceptually relevant losses in training a DCNN for video compression artefact removal. We will use internal representations of a pre-trained classification network as the basis of the loss functions. Specifically, the LPIPS metric and a perceptual discriminator will be responsible for low-level and high-level features respectively. The perceptual discriminator uses differing internal feature representations of the VGG network as its first stage of feature extraction. Initial results shows an increase in performance in perceptually based metrics such VMAF, LPIPS and BRISQUE, while showing a decrease in performance in PSNR.
The volume of User Generated Content (UGC) on the internet has exploded throughout the pandemic. The relatively low quality of that content generally implies an increased bitrate and much reduced quality after transcoding. Preprocessing e.g. using a noise reducer, is one approach for reducing bitrate and increasing quality. The impact of the noise reducer is however affected by the target bitrate of the encoder. That relationship is known but not previously quantitatively examined. In this paper we present a methodology and new metric for measuring this impact based on the Rate-Distortion curves before and after pre-processing. The metric is used as a cost function for estimating the optimal filter parameter for our chosen denoiser. Our experiments show that optimising the filter parameter in this way yields as much as 4-5dB improvement in PSNR at 3 Mbps.
The technology climate for video streaming has vastly changed during 2020. Since the pandemic, video traffic over the internet has increased dramatically. This has clearly put increased interest in the bitrate/quality tradeoff for video compression for applications in video streaming and real time video communications. As far as we know, the impact of different artefacts on that tradeoff has not previously been systematically evaluated. In this paper we propose a methodology for measuring the impact of various degradations (noise, grain, flicker, shake) in a video compression pipeline. We show that noise/grain has the largest impact on codec performance, but that the modern codecs are more robust to the artefact. In addition, we report on the impact of a denoising module deployed as a pre-processor and show that performance metrics change in the context of the pipeline. Denoising would benefit from being treated as part of the processing pipeline both in development and testing.
Traditional metrics for evaluating video quality do not completely capture the nuances of the Human Visual System (HVS), however they are simple to use for quantitatively optimizing parameters in enhancement or restoration. Modern Full-Reference Perceptual Visual Quality Metrics (PVQMs) such as the video multi-method assessment fusion (VMAF) function are more robust than traditional metrics in terms of the HVS, but they are generally complex and non-differentiable. This lack of differentiability means that they cannot be readily used in optimization scenarios for enhancement or restoration. In this paper we look at the formulation of a perceptually motivated restoration framework for video. We deploy this process in the context of denoising by training a spatio-temporal denoiser deep convultional neural network (DCNN). We design DCNNs as a differentiable proxy for both a spatial and temporal version of VMAF. These proxies are used as part of the proposed loss function in updating the weights of the spatio-temporal DCNNs. We use these proxies and traditional losses to propose a perceptually motivated loss function for video. Our results show that using the perceptual loss function as a fine tuning step yields a higher VMAF score and lower PSNR, when compared to the spatio-temporal network that is trained using the traditional mean squared error loss. Using the perceptual loss function for the entirety of training yields a lower VMAF and PSNR, but has visibly less noise in its output.
Modern Perceptual Visual Quality Metrics (PVQMs) for video are generally complex and non-differentiable. This makes them difficult to use as loss functions in restoration and compression tuning. Traditional metrics such as PSNR/MSE which are differentiable remain important but do not capture perceptual visual criteria. In this paper we present a DNN which models a popular perceptual video metric VMAF. In so doing, we introduce a differentiable loss function that closely matches the behaviour of a perceptual metric. Employing degradation generated with H.265 compression, our model achieves a 4.41% RMSE in predicting VMAF. This can now be deployed as a video based loss function in video enhancement and compression tasks.
This paper describes high-quality compression of high dynamic range (HDR) video using existing tools such as the HEVC Main 10 profile, the SMPTE ST 2084 (PQ) transfer function, and the BT.2020 non-constant luminance Y'CbCr color representation. First, we present novel mathematical bounds that reduce complexity of luminance-preserving subsampling (luma adjustment). A nested look-up table allows for further speedup. Second, an adaptive QP scheme is presented that obtains a better bit allocation balance between dark and bright areas of the picture. Third, a method to control the bit allocation balance between chroma and luma by adjusting the chroma QP offset is presented. The result is a considerable increase in perceptual quality compared to the anchors used in the 2015 MPEG High Dynamic Range/Wide Color Gamut Call for Evidence. All techniques are encoder-side-only, making them compatible with a regular decoder capable of supporting HEVC Main10/PQ/BT.2020, which is already available in some TV sets on the market.
A. C. Kokaram合作论文数College Green;Trinity College;Electronic and Electrical Engineering Department5