AV2 is the next-generation video coding standard from the Alliance for Open Media. Relative to AV1, AV2 achieves substantial bitrate savings enabling high-quality video delivery while maintaining practical encoder and decoder complexity. This paper provides a summary of the residual and entropy coding innovations adopted in AV2 including the (i) new coefficient coding design within the adaptive transform coding (ATC) framework, (ii) probability adaptation rate adjustment (PARA) for improved probability estimation in entropy coding, (iii) adaptive truncated Rice coding for compression of large coefficient magnitudes, and (iv) forward skip coding (FSC) for coding transform-skip residuals.
AV2 is the successor to the AV1 video coding standard developed by the Alliance for Open Media (AOMedia). Its primary objective is to deliver substantial compression gains and subjective quality improvements while maintaining low-complexity encoder and decoder operations. This paper describes the transform, quantization and entropy coding design in AV2, including redesigned transform kernels and data-driven transforms, expanded transform partitioning, and a mode coefficient dependent transform signaling. AV2 introduces several new coding tools including Intra/Inter Secondary Transforms (IST), Trellis Coded Quantization (TCQ), Adaptive Transform Coding (ATC), Probability Adaptation Rate Adjustment (PARA), Forward Skip Coding (FSC), Cross Chroma Component Transforms (CCTX), Parity Hiding (PH) tools and improved lossless coding. These advances enable AV2 to deliver the highest quality video experience for video applications at a significantly reduced bitrate.
Recent advances in implicit neural representations (INRs) have shown significant promise in modeling visual signals for various low-vision tasks including image super-resolution (ISR). INR-based ISR methods typically learn continuous representations, providing flexibility for generating high-resolution images at any desired scale from their low-resolution counterparts. However, existing INR-based ISR methods utilize multi-layer perceptrons for parameterization in the network; this does not take account of the hierarchical structure existing in local sampling points and hence constrains the representation capability. In this paper, we propose a new Hierarchical encoding based Implicit Image Function for continuous image super-resolution, HIIF, which leverages a novel hierarchical positional encoding that enhances the local implicit representation, enabling it to capture fine details at multiple scales. Our approach also embeds a multi-head linear attention mechanism within the implicit attention network by taking additional non-local information into account. Our experiments show that, when integrated with different backbone encoders, HIIF outperforms the state-of-the-art continuous image super-resolution methods by up to 0.17dB in PSNR. The source code of HIIF will be made publicly available at https://github.com/YuxuanJJ/HIIF.
Video super-resolution (VSR) is a critical task for enhancing low-bitrate and low-resolution videos, particularly in streaming applications. While numerous solutions have been developed, they often suffer from high computational demands, resulting in low frame rates (FPS) and poor power efficiency, especially on mobile platforms. In this work, we compile different methods to address these challenges, the solutions are end-to-end real-time video super-resolution frameworks optimized for both high performance and low runtime. We also introduce a new test set of high-quality 4K videos to further validate the approaches. The proposed solutions tackle video up-scaling for two applications: 540p to 4K (x4) as a general case, and 360p to 1080p (x3) more tailored towards mobile devices. In both tracks, the solutions have a reduced number of parameters and operations (MACs), allow high FPS, and improve VMAF and PSNR over interpolation baselines. This report gauges some of the most efficient video super-resolution methods to date.
Super-resolution (SR) is a key technique for improving the visual quality of video content by increasing its spatial resolution while reconstructing fine details. SR has been employed in many applications including video streaming, where compressed low-resolution content is typically transmitted to end users and then reconstructed with a higher resolution and enhanced quality. To support real-time playback, it is important to implement fast SR models while preserving reconstruction quality; however, most existing solutions, in particular those based on complex deep neural networks, fail to do so. To address this issue, this paper proposes a low-complexity SR method, RTSR, designed to enhance the visual quality of compressed video content, focusing on resolution up-scaling from a) 360p to 1080p and from b) 540p to 4K. The proposed approach utilizes a Convolutional Neural Network (CNN)-based network architecture, which was optimized for AOMedia Video 1 (AV1-SVT)-encoded content at various quantization levels based on a dual-teacher knowledge distillation method. This method was submitted to the AIM 2024 Video Super-Resolution Challenge, specifically targeting the Efficient/Mobile Real-Time Video Super-Resolution competition. It achieved the best trade-off between complexity and coding performance (measured in PSNR, SSIM and VMAF) among all six submissions. The code will be available at https://github.com/YuxuanJJ/RTSR.
In recent years, attention mechanisms have been exploited in single image super-resolution (SISR), achieving impressive reconstruction results. However, these advancements are still limited by the reliance on simple training strategies and network architectures designed for discrete up-sampling scales, which hinder the model's ability to effectively capture information across multiple scales. To address these limitations, we propose a novel framework, C2D-ISR, for optimizing attention-based image super-resolution models from both performance and complexity perspectives. Our approach is based on a two-stage training methodology and a hierarchical encoding mechanism. The new training methodology involves continuous-scale training for discrete scale models, enabling the learning of inter-scale correlations and multi-scale feature representation. In addition, we generalize the hierarchical encoding mechanism with existing attention-based network structures, which can achieve improved spatial feature fusion, cross-scale information aggregation, and more impor-tantly, much faster inference. We have evaluated the C2D-ISR framework based on three efficient attention-based backbones, SwinIR-L, SRFormer-L and MambaIRv2-L, and demonstrated significant improvements over the other ex-isting optimization framework, HiT, in terms of super-resolution performance (up to 0.2dB) and computational complexity reduction (up to 11%).
Deep learning is now playing an important role in enhancing the performance of conventional hybrid video codecs. These learning-based methods typically require diverse and representative training material for optimization in order to achieve model generalization and optimal coding performance. However, existing datasets either offer limited content variability or come with restricted licensing terms constraining their use to research purposes only. To address these issues, we propose a new training dataset, named BVI-AOM, which contains 956 uncompressed sequences at various resolutions from 270p to 2160p, covering a wide range of content and texture types. The dataset comes with more flexible licensing terms and offers competitive performance when used as a training set for optimizing deep video coding tools. The experimental results demonstrate that when used as a training set to optimize two popular network architectures for two different coding tools, the proposed dataset leads to additional bitrate savings of up to 0.29 and 2.98 percentage points in terms of PSNR-Y and VMAF, respectively, compared to an existing training dataset, BVI-DVC, which has been widely used for deep video coding. The BVI-AOM dataset is available at https://github.com/fan-aaron-zhang/bvi-aom.
Banding is a visually unpleasing artifact appearing in flat areas of encoded content that no video standard has fully addressed. We propose a normative debanding filter to tackle banding artifacts and have tested it as an in-loop and post-loop filter in AVM. Debanding is achieved by introducing dithering on a frame level to the luma component. The proposed filter shows CAMBI gains for content with banding while not affecting other content. Although the added dithering has a minor negative impact on some objective metrics, subjective improvements in banding-prone content are (informally) observed. On the test set, encoding time increases on average by 0.5%, while decoding time increases by around 0.5% for in-loop and 1.5% for post-loop.
Banding is manifested as false contours in otherwise smooth regions in an image or a video. There are many different reasons for banding artifacts to occur, however, one of the most prominent causes is the quantization inside a video encoder. Compared to other types of artifacts common in video processing, e.g. blur, ringing, or blockiness, only a relatively small change of the original pixel values can produce an easily noticeable and visually very annoying case of banding. This property makes it very difficult for banding to be captured by generic objective quality metrics such as PSNR or VMAF [10] which brings a need for a distortion specific detector targeted directly at banding artifacts. Most of the previous attempts to solve this problem tried to tackle it as false segments or false edges detection. Both block-based [7, 14] and pixel-based [2, 3, 15] segmentation methods have been tried in the first category, while the edge-based methods exploited different local statistics such as gradients, contrast, or entropy [4, 6, 9, 13]. The main difficulty for all of these approaches is distinguishing between the real and false edges or segments. Recently, banding detection has also been addressed by deep neural networks [8]. The above mentioned approaches have been developed for 8-bit content and mostly tuned towards the banding artifacts occurring in user-generated images and videos. Moreover, they do not address the potential presence of dithering - an intentionally inserted noise used to randomize the error caused by quantization. Dithering is commonly used during bit-depth conversion and is often enabled by default in popular image and video processing tools, such as ffmpeg [12]. Despite being highly effective in reducing the perceived banding, dithering does not suppress the false contours completely, and thus needs to be factored in a reliable banding detector. Our goal was, therefore, to develop an algorithm capable of evaluating perceived banding in professionally generated videos processed in ways relevant to the adaptive streaming scenario (i.e. video compression and scaling). The requirements also included ability to capture the effect of dithering and to work on both 8-bit and 10-bit content. We hereby present CAMBI, a Contrast Aware Multiscale Banding Index. CAMBI is a white-box solution to the above described problem derived from basic principles of human vision with just a few, perceptually-motivated, parameters. The first version was introduced at PCS'2021 [11]. Here, we also present several improvements made since then. There are three main steps in CAMBI - input preprocessing, multiscale banding confidence calculation, and spatio-temporal pooling. Although it has been shown that chromatic banding exists [5], like most past works, we assume that most of the banding can be captured in the luma channel. The preprocessing step, therefore, consists of luma channel extraction followed by filtering to account for dithering and a spatial mask computation to exclude regions with textures. Banding confidence is calculated for 4 brightness level differences on 5 scales, taking into account contrast perception of human visual system. This creates 20 banding confidence maps per frame that are pooled spatially considering only a certain percentage of highest banding confidence. Such mechanism ensures that even the banding appearing in relatively small area of the frame is captured proportionally to its perceptual importance. Finally, the scores from different video frames are pooled into a single banding index. To test the accuracy of CAMBI, we conducted a subjective test on 86 video clips created from 9 different sources from Netflix catalog using different levels of compression and scaling with and without dithering. The ground-truth mean opinion scores (MOS) were obtained from 26 observers who were asked to rate the annoyance of the banding in the scene on the continuous impairment scale annotated with 5 equidistant labels (imperceptible, perceptible but not annoying, slightly annoying, annoying, very annoying) [1]. CAMBI achieved a correlation exceeding 0.94 in terms of both Pearson Linear Correlation Coefficient (PLCC) and Spearman Rank Order Correlation Coefficient (SROCC), significantly outperforming state-of-the-art banding detectors for our use-case. CAMBI is currently used alongside VMAF in our production to improve the quality of encodes prone to banding. In the future, we are planning to integrate it into VMAF as one of the features to make it capable of accurately evaluating video quality in the presence of banding as well as other artifacts.
Staircase-like contours introduced to a video by quantization in flat areas, commonly known as banding, have been a longstanding problem in both video processing and quality assessment communities. The fact that even a relatively small change of the original pixel values can result in a strong impact on perceived quality makes banding especially difficult to be detected by objective quality metrics. In this paper, we study how banding annoyance compares to more commonly studied scaling and compression artifacts with respect to the overall perceptual quality. We further propose a simple combination of VMAF and the recently developed banding index, CAMBI, into a banding-aware video quality metric showing improved correlation with overall perceived quality.
In HTTP Adaptive Streaming, video content is conventionally encoded by adapting its spatial resolution and quantization level to best match the prevailing network state and display characteristics. It is well known that the traditional solution, of using a fixed bitrate ladder, does not result in the highest quality of experience for the user. Hence, in this paper, we introduce a content-driven approach for estimating the bitrate ladder, based on spatio-temporal features extracted from the uncompressed content. The method implements a content-driven interpolation. It uses the extracted features to train a machine learning model to infer the curvature points of the Rate-VMAF curves in order to guide a set of initial encodings. We employ the VMAF quality metric as a means of perceptually conditioning the estimation. When compared to the generation of a reference ladder using exhaustive encoding, 76.63% the estimated ladder's Rate-VMAF points are identical to those of the reference ladder. The proposed method benefits from a significant (77.4%) reduction in the number of encodes required with only a small (1.04%) average Bj⊘ntegaard Delta Rate increase.
One of the challenges faced by many video providers is the heterogeneity of network specifications, user requirements, and content compression performance. The universal solution of a fixed bitrate ladder is inadequate in ensuring a high quality of user experience without re-buffering or introducing annoying compression artifacts. However, a content-tailored solution, based on extensively encoding across all resolutions and over a wide quality range is highly expensive in terms of computational, financial, and energy costs. Inspired by this, we propose an approach that exploits machine learning to predict a content-optimized bitrate ladder for on-demand video services. The method extracts spatio-temporal features from the uncompressed content, trains machine-learning models to predict the Pareto front parameters and, based on that, builds the ladder within a defined bitrate range. The method has the benefit of significantly reducing the number of encodes required per sequence. The presented results, based on 100 HEVC-encoded sequences, demonstrate a reduction in the number of encodes required when compared to an exhaustive search and an interpolation-based method, by 89.06% and 61.46%, respectively, at the cost of an average Bjøntegaard Delta Rate difference of 1.78% compared to the exhaustive approach. Finally, a hybrid method is introduced that selects either the proposed or the interpolation-based method depending on the sequence features. This results in an overall 83.83% reduction of required encodings at the cost of an average Bjøntegaard Delta Rate difference of 1.26%.
In many image and video processing applications, the ability to resize by a fractional factor, such as from 1080p to 720p, is essential. However, conventional CNN layers can only be used to alter the resolution of their inputs with integer scale factors. In this paper, we propose a downsampling network architecture that progressively reconstructs residuals at different scales. In particular, the aforementioned problem is solved by combining an upsampling sub-network and a downsampling subnetwork, both with integer scale factor. As an application, we apply the proposed downsampling network to an adaptive bitrate video streaming scenario. We extensively evaluate with different video codecs and upsampling algorithms to show the generality of our model. Our experimental results show that improvements in coding efficiency over the conventional Lanczos downsampling and state-of-the-art methods are attained, measured in different perceptual video quality models on large-resolution test videos.
Banding artifacts are artificially-introduced contours arising from the quantization of a smooth region in a video. Despite the advent of recent higher quality video systems with more efficient codecs, these artifacts remain conspicuous, especially on larger displays. In this work, a comprehensive subjective study is performed to understand the dependence of the banding visibility on encoding parameters and dithering. We subsequently develop a simple and intuitive no-reference banding index called CAMBI (Contrast-aware Multiscale Banding Index) which uses insights from Contrast Sensitivity Function in the Human Visual System to predict banding visibility. CAMBI correlates well with subjective perception of banding while using only a few visually-motivated hyperparameters.
Video coding in the YCbCr color space has been widely used, since it is efficient for compression, but it can result in color distortion due to conversion error. Meanwhile, coding in the RGB color space maintains high color fidelity, having the drawback of a substantial bitrate increase with respect to YCbCr coding. Cross-component prediction (CCP) efficiently compresses video content by decorrelating color components while keeping high color fidelity. In this scheme, the chroma residual signal is predicted from the luma residual signal inside the coding loop. This paper gives a description of the CCP scheme from several points of view, from theoretical background to practical implementation. The proposed CCP scheme has been evaluated in standardization communities and adopted into H.265/High Efficiency Video Coding (HEVC) Range Extensions. The experimental results show significant coding performance improvements for both natural and screen content video, while the quality of all color components is maintained. The average coding gains for natural video are 17% and 5% bitrate reduction in the case of intra coding and 11% and 4% in the case of inter coding for RGB and YCbCr coding, respectively, while the average increment of encoding and decoding times in the HEVC reference software implementation are 10% and 4%, respectively.
Measuring the quality of digital videos viewed by human observers has become a common practice in numerous multimedia applications, such as adaptive video streaming, quality monitoring, and other digital TV applications. Here we explore a significant, yet relatively unexplored problem: measuring perceptual quality on videos arising from both luma and chroma distortions from compression. Toward investigating this problem, it is important to understand the kinds of chroma distortions that arise, how they relate to luma compression distortions, and how they can affect perceived quality. We designed and carried out a subjective experiment to measure subjective video quality on both luma and chroma distortions, introduced both in isolation as well as together. Specifically, the new subjective dataset comprises a total of 210 videos afflicted by distortions caused by varying levels of luma quantization commingled with different amounts of chroma quantization. The subjective scores were evaluated by 34 subjects in a controlled environmental setting. Using the newly collected subjective data, we were able to demonstrate important shortcomings of existing video quality models, especially in regards to chroma distortions. Further, we designed an objective video quality model which builds on existing video quality algorithms, by considering the fidelity of chroma channels in a principled way. We also found that this quality analysis implies that there is room for reducing bitrate consumption in modern video codecs by creatively increasing the compression factor on chroma channels. We believe that this work will both encourage further research in this direction, as well as advance progress on the ultimate goal of jointly optimizing luma and chroma compression in modern video encoders.
A challenge that many video providers face is the heterogeneity of networks and display devices for streaming, as well as dealing with a wide variety of content with different encoding performance. In the past, a fixed bit rate ladder solution based on a "fitting all" approach has been employed. However, such a content-tailored solution is highly demanding; the computational and financial cost of constructing the convex hull per video by encoding at all resolutions and quantization levels is huge. In this paper, we propose a content-gnostic approach that exploits machine learning to predict the bit rate ranges for different resolutions. This has the advantage of significantly reducing the number of encodes required. The first results, based on over 100 HEVC-encoded sequences demonstrate the potential, showing an average Bjøntegaard Delta Rate (BDRate) loss of 0.51% and an average BDPSNR loss of 0.01 dB compared to the ground truth, while significantly reducing the number of pre-encodes required when compared to two other methods (by 81%-94%).
Palette mode is a new coding tool included in the HEVC screen content coding extension (SCC) to improve the coding efficiency for screen contents such as computer generated video with substantial amount of text and graphics. It is observed that a local area in screen content typically has a few colors separated by sharp edges. To exploit such characteristics, palette mode represents samples in a block with indexes pointing to the color entries in a palette table. This paper provides a detailed overview of the palette mode in HEVC SCC in terms of palette generation, coding of the palette, and coding of the palette indexes for the samples in the palette block. Several improvements to palette mode coding, which have been proposed but not included in HEVC SCC, are also described. Simulation results are presented to quantify the bitrate savings provided by the palette mode for equal distortions.
The Range Extensions (RExt) of the High Efficiency Video Coding (HEVC) standard have recently been approved by both ITU-T and ISO/IEC. This set of extensions targets video coding applications in areas including content acquisition, postproduction, contribution, distribution, archiving, medical imaging, still imaging, and screen content. In addition to the functionality of HEVC Version 1, RExt provide support for monochrome, 4: 2: 2, and 4: 4: 4 chroma sampling formats as well as increased sample bit depths beyond 10 bits per sample. This extended functionality includes new coding tools with a view to provide additional coding efficiency, greater flexibility, and throughput at high bit depths/rates. Improved lossless, near-lossless, and very high bit-rate coding is also a part of the RExt scope. This paper presents the technical aspects of HEVC RExt, including a discussion of RExt profiles, tools, applications, and provides experimental results for a performance comparison with previous relevant coding technology. When compared with the High 4: 4: 4 Predictive Profile of H.264/Advanced Video Coding (AVC), the corresponding HEVC 4: 4: 4 RExt profile provides up to similar to 25%, similar to 32%, and similar to 36% average bit-rate reduction at the same PSNR quality level for intra, random access, and low delay configurations, respectively.
This paper presents a method for efficient compression of high dynamic range (HDR) and wide color gamut (WCG) video data. The proposed solution consists of two major elements: a conventional video codec (e.g., HEVC) and pre-and post-processing steps applied prior to encoding and after decoding process, respectively. The proposed HDR/WCG video coding system can be configured to provide two configurations: (1) a non-backward compatible bitstream with improved HDR video quality and (2) a SDR backward compatible bitstream with balanced visual quality between the reconstructed signal by the SDR and the HDR receivers. The simulations conducted under the MPEG Common Test Conditions for HDR demonstrate that the compression efficiency of the proposed solution outperforms the anchor solution on objective metrics. Additionally, subjective evaluations conducted under MPEG revealed improved visual quality for the proposed method.