Augmented reality (AR) and virtual reality (VR) applications have seen rising popularity in recent years. Omnidirectional 360° video is a video format often used in AR and VR applications. To address the industry needs, a new HEVC edition recently published includes several supplemental enhancement information (SEI) messages to enable the carriage of omnidirectional video using HEVC. However, further improvement in 360° video compression efficiency is needed. In order to address this challenge, the Joint Video Exploration Team (JVET) of ITU-T VCEG and ISO/IEC MPEG has been investigating 360° video coding technologies, including projection formats, pre- and post-processing technologies, as well as 360°-video-specific coding tools since 2016. The joint call for proposals (CfP) recently issued by ITU-T VCEG and ISO/IEC MPEG on video compression technologies beyond HEVC included a category on 360° video. Twelve CfP responses in the 360° video category were received. This paper describes technologies relevant to 360° video for VVC. A summary of projection formats, pre- and post-processing methods, and 360°-video specific coding tool modifications in these proposals is provided.
The ITU-T Video Coding Experts Group (VCEG) and ISO/IEC Moving Picture Experts Group (MPEG) issued in October 2017 a joint Call for Proposals (CfP) on video compression with capability beyond HEVC. The joint CfP included three categories of content: standard dynamic range (SDR), high dynamic range and wide color gamut (HDR/WCG), and 360 degrees omni-directional video (360 degrees). This paper describes a response to the joint CfP that considers all three categories of video content. The core codec in the response is designed based on the joint exploration model (JEM) reference software. The key coding tools in the JEM are significantly simplified to reduce both average and worst-case complexity for hardware design with negligible coding performance loss. Furthermore, two additional coding tools are used to further improve coding efficiency. For the HDR and 360 degrees categories, additional coding tools specifically designed to optimize the compression efficiency and subjective quality of that specific content category are included. Further, some SDR coding tools are modified to alleviate subjective quality problems. For the random access configuration, compared to the HEVC test model (HM) anchor, the proposed video codec achieves average luma rate savings of 35.7%, 31.3%, and 33.9% for the SDR, HDR, and 360 degrees categories, respectively.
Standards play an important role in providing a common set of specifications and allowing inter-operability between devices and systems. Until recently, no standard for high-dynamic-range (HDR) image coding had been adopted by the market, and HDR imaging relies on proprietary and vendor-specific formats which are unsuitable for storage or exchange of such images. To resolve this situation, the JPEG Committee is developing a new coding standard called JPEG XT that is backward compatible to the popular JPEG compression, allowing it to be implemented using standard 8-bit JPEG coding hardware or software. In this paper, we present design principles and technical details of JPEG XT. It is based on a two-layer design, a base layer containing a low-dynamic-range image accessible to legacy implementations, and an extension layer providing the full dynamic range. The paper introduces three of currently defined profiles in JPEG XT, each constraining the common decoder architecture to a subset of allowable configurations. We assess the coding efficiency of each profile extensively through subjective assessments, using 24 naïve subjects to evaluate 20 images, and objective evaluations, using 106 images with five different tone-mapping operators and at 100 different bit rates. The objective results (based on benchmarking with subjective scores) demonstrate that JPEG XT can encode HDR images at bit rates varying from 1.1 to 1.9 bit/pixel for estimated mean opinion score (MOS) values above 4.5 out of 5, which is considered as fully transparent in many applications. This corresponds to 23-times bitstream reduction compared to lossless OpenEXR PIZ compression.
This paper describes a video coding scheme submitted in response to the joint call for proposals (CfP) on video compression for capability beyond HEVC issued by ITU-T SG16 Q.6 (VCEG) and ISO/IEC JTC1/SC29/WG11(MPEG) in October 2017. It includes video coding techniques for the standard dynamic range (SDR) and high dynamic range (HDR) categories. Design of the core SDR codec in the response is based on the joint exploration model (JEM) reference software developed by the joint video exploration team (JVET). Some of key coding tools in the JEM are significantly simplified to reduce both average and worst-case complexity for hardware design with negligible coding performance loss. Furthermore, two additional coding technologies, namely multi-type tree (MTT) and decoder-side intra mode derivation (DIMD), are used to further improve coding efficiency. For the HDR category, besides the tools used in SDR category, two additional coding tools: an in-loop reshaper and a luma-based QP prediction method are used to further improve HDR coding efficiency. Simulation results demonstrate the high coding efficiency achieved by the proposed video codec at the expense of moderate coding complexity over HEVC. For random access configuration, it achieves average bit rate savings of 35.7% and 4.00% over the HM and JEM anchors with decoding time of 263% and 33%, respectively, for the SDR sequences. For the HDR sequences, the proposed in-loop reshaper is configured to maximize HDR objective metrics, it achieves average bit rate savings of 31.3% and 4.6% over the HM and JEM for wPSNRY metrics for the HDR PQ content.
360-degree video is emerging as a new way of offering immersive visual experience. The quality evaluation of 360-degree video is more difficult compared to the quality evaluation of conventional video. However, to ensure successful development of 360-degree video coding technologies, it is essential to precisely measure both objective and subjective quality. In this paper, an overview of the 360-degree video quality evaluation framework established by the joint video exploration team (JVET) of ITU-T VCEG and ISO/IEC MPEG is provided. This framework aims at reproducing the different processes in the 360-degree video processing workflow that are related to coding. The results of different experiments conducted using the JVET framework are reported to illustrate the impact on objective and subjective quality with different projection formats and codecs.
This paper describes a 360 degrees video coding scheme submitted in response to the joint call for proposal on video compression for capability beyond HEVC issued by ITU-T SG16 Q.6 (VCEG) and ISO/IEC JTCUSC29/WG11 (MPEG) in October 2017. The proposed coding scheme uses projection format adaptation and spherical neighboring relationship to improve coding efficiency. A hybrid angular cubemap projection format is used to adapt the sampling within each face. The faces are packed in an adaptive manner, and the frame packing configuration is updated at regular intervals. Geometry padding of reference photographs is used to improve inter prediction. For intra-/inter-prediction and in-loop filters, processes are modified to avoid utilizing "wrong" neighbors in the frame packed photograph. Finally, post-filtering is used to reduce artifacts at face discontinuities. Experimental results are presented to demonstrate the superior compression efficiency achieved by the proposed 360 degrees video coding scheme. Based on the end-to-end weighted-to-spherically uniform PSNR metric, it achieves average bit rate savings of 33.9% and 13.5% over the HM and joint exploration model anchors, respectively, encoded in the padded equirectangular projection format.
In this paper, a novel hybrid cubemap projection (HCP) is proposed to improve the 360-degree video coding efficiency. HCP allows adaptive sampling adjustments in the horizontal and vertical directions within each cube face. HCP parameters of each cube face can be adjusted based on the input 360-degree video content characteristics for a better sampling efficiency. The HCP parameters can be updated periodically to adapt to temporal content variation. An efficient HCP parameter estimation algorithm is proposed to reduce the computational complexity of parameter estimation. Experimental results demonstrate that HCP format achieves on average luma (Y) BD-rate reduction of 11.51%, 8.0%, and 0.54% compared to equirectangular projection format, cubemap projection format, and adjusted cubemap projection format, respectively, in terms of end-to-end WS-PSNR.
360-degree video has become popular in recent years with the advances in virtual reality (VR) and augmented reality (AR) technologies and has been rapidly commercialized. To provide viewers with an immersive experience, 360-degree video requires higher resolution and much higher bandwidth compared with conventional 2D video. In a typical 360-degree video compression and delivery framework, the stitched input 360-degree videos, represented in a native projection format, e.g., equirectangular (ERP), are converted into another projection format, e.g., cubemap (CMP), octahedron (OHP), etc. and frame packed before being fed into existing video codecs. The intermediate projection format is important and would potentially improve the representation efficiency and coding performance. Among all the projection solutions, CMP is very popular and has been widely used in the computer graphics community. The intrinsic rectilinear properties of the CMP format are advantageous for the translational motion model in the modern codec architecture. However, in the CMP representation, the samples on the sphere are not evenly distributed within the faces, resulting in a higher density near the face boundaries and a lower density near the face center. Such non-uniform sampling scheme penalizes the video representation efficiency and degrades the coding performance. Adjusted cubemap projection (ACP) was proposed to address such non-uniform sampling by introducing transform functions to improve the sampling uniformity. However, the transform function parameters in ACP are fixed regardless of the content inside each cube face. In this paper, a generalized hybrid cubemap projection (HCP) is proposed to improve the 360-degree video coding efficiency beyond ACP. HCP is defined by a pair of forward transform and inverse transform functions with a pair of horizontal and vertical transform parameters per cube face. The encoder can choose the optimal sampling for each face by adjusting the parameters in the horizontal and vertical directions based on the 360-degree video content characteristics inside each cube face. In order to maintain the boundary continuities between two neighboring faces, in a 3x2 packing layout, vertical parameter constraints are imposed such that faces in each face-row have the same vertical parameters. The HCP parameters are chosen to minimize the end-to-end weighted conversion error and determined using iterative search between the horizontal and the vertical directions. Significant changes in HCP parameter values can cause drastic change in sampling distribution, and may affect the inter-picture coding efficiency. Therefore, an efficient HCP parameter estimation algorithm is proposed to achieve a better trade-off between the temporal sampling adaptation and the inter-picture prediction efficiency by reducing the temporal variation of HCP parameters. The proposed HCP parameter search algorithm reduces the computational complexity by 5x compared to the exhaustive search method. The HCP parameters are selected by the encoder using the first picture of each Intra Random-Access Point (IRAP) and signalled once per IRAP. In SPS, projection format, frame packing parameters including number of faces in horizontal and vertical directions and each face's position and orientation are signalled. In PPS, the horizontal and vertical HCP parameters in 6-bit precision are encapsulated. The proposed HCP solution is implemented upon JEM-6.0 and 360Lib-3.0 software. Simulation results are reported using the test conditions specified in the JVET Call-for-Evidence (CfE) document. Compared with the CMP and ACP formats, the proposed HCP format demonstrates average 3.0 dB (up to 3.6 dB) and 0.2dB (up to 0.4 dB) End-to-End WS-PSNR improvement for the luma (Y) component, respectively, and average luma (Y) BD-rate reductions of 11.5% (up to 23.0%) and 0.5% (up to 1.0%), respectively.
Modern cameras can be used to capture high dynamic range (HDR) images by taking multiple low dynamic range (LDR) pictures at different exposure settings. However, DSLR cameras often store the multiple exposure pictures, which increases storage requirements proportionally to the number of exposures, while smartphones and tablets only store a tone-mapped representation of the HDR image instead of the original HDR data. In this paper, we investigate the bit rate reduction that could be achieved by storing HDR images compressed with JPEG XT, a JPEG backward-compatible HDR image compression standard, over multiple exposures JPEG. We also investigate the quality increase that could be achieved by fusing images with linear values and native sensor bit depth over pictures with sRGB gamma encoded values and 8 bits per sample. To investigate the impact of the different parameters on the performance, we perform a set of experiments with 14 contents, 4 resolutions (ranging from 2.5 megapixels (MP) to 16MP), 4 exposures configurations (2, 3, 5, and 7 exposures), 5 tone mapping operators (TMOs) used to generate the JPEG XT base layer, and 5 quality levels representative of values commonly used in smartphones and DSLR cameras. Results show that the performance is dependent on the TMO and number of exposures, with maximum average bit rate savings ranging from 10% for two exposures up to 70% for seven exposures. Additionally, quality improvements of 1.7 dB to 5.1 dB can be achieved when using the native image sensor data format.
The paper presents a new geometry padding method for motion compensated prediction in 360 video coding. Unlike the conventional padding method for 2D video coding, which extends samples outside of a picture's boundary by simply copying (repeating) those at the boundaries, the proposed geometry padding method considers the spherical nature of the 360 video and the specific geometry projection format when extending samples outside of a picture's or a face's boundary. The proposed geometry padding method is implemented in the HEVC reference software HM-16.12 for the equirectangular and cubemap projection formats. Simulation results show that, when compared with the conventional 2D padding method in HEVC, geometry padding performs better; the proposed geometry padding gives on average luma (Y) BD-rate reduction of 0.3% for equirectangular projection format coding and 1.0% for cubemap projection format coding all in terms of the spherical PSNR (S-PSNR) metric. The proposed geometry padding is especially effective for high motion sequences, where up to 4.3% BD-rate reduction can be achieved.
360 Video has become popular in recent years, as commercial interests in deploying Virtual Reality (VR) applications rise. This type of video is usually captured using multi-camera arrays, such as the GoPro Omni camera rig. After separate video streams are captured from multiple cameras, image stitching is applied to obtain a spherical representation of the scene, which spans 360 degrees horizontally and 180 degrees vertically, hence the name 360 video. In the existing workflow of 360 spherical video coding, the 360 video is projected onto the 2D plane with a projection format, such as equirectangular (ERP), cubemap (CMP), equal-area (EAP), octahedron (OHP), etc. Most, if not all, of the currently available 360 video content are provided in ERP format defined in longitude and latitude. Projection format conversion may be performed to convert the native ERP format to another format before coding is applied. Some projection formats contain more than one face, for example, CMP projects the sphere onto a cube of six faces or OHP projects the sphere onto an octahedron of eight faces. For these multi-face projection formats, the faces are packed onto a 2D rectangular picture with a frame packing method. For example, the six faces of CMP can be packed with 4x3 configuration, or 3x2 configuration. Finally, the frame packed picture is coded as a 2D conventional video. Existing video codecs are designed only considering conventional 2D video captured on a plane. When motion compensated prediction uses any samples outside of a reference picture's boundaries, padding will be performed by simply copying the sample values from the picture boundaries. This repetitive padding method is referred as conventional 2D padding method, which is widely used in video coding standards such as H.264, High Efficiency Video Coding (HEVC). However, a 360 video encompasses video information on the whole sphere, and thus intrinsically has a cyclic property. When considering this cyclic property, the reference pictures of a 360 video no longer have "boundaries", as the information they contain is all wrapped around a sphere. This cyclic property holds regardless of which projection format or which frame packing is used to represent the 360 video on a 2D plane. The paper presents a new geometry padding method for motion compensated prediction in 360 video coding. Unlike the conventional padding method for 2D video coding, the proposed geometry padding method extends samples outside of a 2D picture's boundaries using neighboring samples on the sphere. The geometry projection format is considered when performing padding. The corresponding sample outside of a face's boundary (which may come from another side in the same face or from another face), is derived with rectilinear projection. Each face is extended with geometry padding separately. When visualized, the extended faces using geometry padding show continuous texture representing natural extension of the texture inside the face. The proposed geometry padding method is implemented in the HEVC reference software HM-16.12 for the ERP and CMP projection formats. In the simulation, a total of sixteen 4K ERP video and eight 8K ERP video are used. For 8K ground truth 8K video, they are converted to 4K video in ERP and CMP projection formats, coded, and converted back to reconstructed 8K video in ERP format. For 4K ground truth video, they are directly coded in 4K ERP, or converted to CMP consisting of 75% of effective samples, coded, and converted back to reconstructed 4K video in ERP format. Then, the end-to-end spherical PSNR (S-PSNR) is calculated between the original 8K or 4K and the reconstructed 8K or 4K ERP video. BD-rate is calculated between the reference unmodified HEVC, which uses the conventional padding method, and HEVC modified with the proposed geometry padding method. Simulation results showed that geometry padding performs better. For 8K sequences, the proposed geometry padding gives on average luma (Y) BD-rate reduction of 0.3% for ERP and 0.8% for CMP, for 4K sequences, the proposed geometry padding gives on average Y BD-rate reduction of 0.2% for ERP and 1.0% for CMP. Comparing the gains in ERP format with the gains in CMP format, the improvement for CMP is larger. This is because CMP has six faces, therefore the improvement from geometry padding method affects more out-of-boundary samples. The proposed geometry padding method is also especially effective for sequences with fast motion. For example, it achieves BD rate reductions of 4.3%, 2.7%, 1.9%, and 2.5% for Glacier, Chairlift, Sb_in_lot, and Driving, respectively. These four sequences are all captured using moving cameras and have fast moving objects. As a result, the sequences contain a lot of across-the-face-boundary motion which can benefit from improved padding method. Detailed simulation results can be found in JVET contribution JVET-D0075 available at http://phenix.int-evry.fr/jvet/doc_end_user/documents/4_Chengdu/wg11/JVET-D0075-v3.zip.
This paper presents a new reference sample derivation method for intra prediction in 360-degree video coding. Unlike the conventional reference sample derivation method for 2D video coding, which uses the samples located directly above and on the left of the current block, the proposed method considers the spherical nature of 360-degree video when deriving reference samples located outside the current face to which the block belongs, and derives reference samples that are geometric neighbors on the sphere. The proposed reference sample derivation method was implemented in the Joint Exploration Model 3.0 (JEM-3.0) for the cubemap projection format. Simulation results for the all intra configuration show that, when compared with the conventional reference sample derivation method, the proposed method gives, on average, luma BD-rate reduction of 0.3% in terms of the weighted spherical PSNR (WS-PSNR) and spherical PSNR (S-PSNR) metrics.
Evaluation methodologies provide a better understanding of the relationship between a technique and the image attributes. Metrics are used to evaluate the similarities between images. They may use different approaches depending on what needs to be achieved. If only objective values need to be compared statistics-based metrics are suitable. A number of image comparison metrics have been proposed in the literature that is based mostly on images statistics. This chapter shows the MATLAB code for the root mean square error (RMSE) calculation. It presents the MATLAB code for the computation of mean square error (MSE) between two images. Peak signal to noise ratio (PSNR) is another widely used metric, which takes into account the maximum value of the signal, and can be defined based on MSE. Metrics enable automation and the use of metrics can also be applied directly to methods, for example, to …
High bit depth data acquisition and manipulation have been largely studied at the academic level over the last 15 years and are rapidly attracting interest at the industrial level. An example of the increasing interest for high-dynamic range (HDR) imaging is the use of 32-bit floating point data for video and image acquisition and manipulation that allows a variety of visual effects that closely mimic the real-world visual experience of the end user [1] (see Figure 1). At the industrial level, we are witnessing increasing traction toward supporting HDR and wide color gamut (WCG). WCG leverages HDR for each color channel to display a wider range of colors. Consumer cameras are currently available with a 14- or 16-bit analog-to-digital converter. Rendering devices are also appearing with the capability to display HDR images and video with a peak brightness of up to 4,000 nits and to support WCG (ITU-R Rec. BT.2020 [2]) rather than the historical ITU-R Rec. BT.709 [3]. This trend calls for a widely accepted standard for higher bit depth support that can be seamlessly integrated into existing products and applications. While standard formats such as the Joint Photographic Experts Group (JPEG) 2000 [5] and JPEG XR [6] offer support for high bit depth image representations, their adoption requires a nonnegligible investment that may not always be affordable in existing imaging ecosystems, and induces a difficult transition, as they are not backward-compatible with the popular JPEG image format.
The procedures commonly used to evaluate the performance of objective quality metrics rely on ground truth mean opinion scores and associated confidence intervals, which are usually obtained via direct scaling methods. However, indirect scaling methods, such as the paired comparison method, can also be used to collect ground truth preference scores. Indirect scaling methods have a higher discriminatory power and are gaining popularity, for example in crowdsourcing evaluations. In this paper, we present how the classification errors, an existing analysis tool, can also be used with subjective preference scores. Additionally, we propose a new analysis tool based on the receiver operating characteristic analysis. This tool can be used to further assess the performance of objective metrics based on ground truth preference scores. We provide a MATLAB script with an implementation of the proposed tools and we show one example of application of the proposed tools.
This paper reports the details and results of a subjective and objective quality evaluation assessing responses to an MPEG call for evidence (CfE) on high dynamic range (HDR) and wide color gamut video coding. Five HDR video contents, compressed at four bit rates by each proponent responding to the CfE, were used in the subjective assessments. To be able to evaluate the performance of objective quality metrics, the double stimulus impairment scale (DSIS) method was used for subjective assessments instead of previously published paired comparison to an anchor. Subjective results show evidence that coding efficiency can be improved in a statistically noticeable way over the HEVC anchor in terms of perceived quality. However, when compared to paired comparison, less statistically significant differences are observed because of the lower discrimination power of the DSIS method. The collected subjective scores were used as a ground truth to benchmark and analyze the performance of objective metrics. Results show that HDR-VDP-2 and PQ2VIFP have the highest correlation with subjective scores and outperform other investigated metrics.
The main objective of this paper is to verify test methodologies for the assessment of high dynamic range (HDR) video.To achieve this, a next generation HDR monitor by Dolby Laboratories was used to display professionally produced HDR content.Two complementary approaches for subjective assessment of HDR video were then designed and carried out at the European Broadcast Union and Ecole Polytechnique Fédérale de Lausanne premises.Results obtained from both evaluations were highly correlated, which shows that they offer a good degree of reliability and reproducibility in different situations.Analysis of the scores in both cases also shows good confidence intervals for each point under test.Finally, they could demonstrate that an increase in terms of quality of experience can be expected from the conventional level of 100 nits to HDR/high brightness at 4000 nits, with intermediate improvements at 400 and 1000 nits.
Scott J. Daly合作论文数Dolby Laboratories2