Traditional subjective quality assessment protocols are designed to produce global quality scores. However, in some use cases such as watermarking, there is a need for local subjective assessment since the traditional methods lack the granularity required to generate precise local quality maps. In this paper, we propose LISA (Local Impairment Scale Annotator), a new subjective protocol and supporting annotation tool that enable pixel-wise local quality assessment. We present the complete LISA protocol specifications, from the temporal organization of the test to the user interface design. Finally, we illustrate LISA's deployment in a use case scenario evaluating highfidelity digital watermarking. The annotation tool is available at https://github.com/edemezet-nagra/LISA-subjective-protocol
3D Gaussian Splatting (3DGS) is a major reference for learning 3D models of real scenes, enabling real-time novel view synthesis and fast model training. As part of the Elliptical Weighted Average (EWA) volume splatting methods, 3DGS assumes that the support areas of the Gaussians do not overlap, which – among other hypothesis – leads to a simple rendering by alpha blending. While this algorithm enables training a 3D model, theoretical analysis shows that alpha blending generates rendering misalignments for Gaussians belonging to the same surface, which may impair learning and eventually view synthesis. In this paper, we analyze the difference between alpha blending of Gaussian kernels and an accumulation by weighted average of their contributions, and we propose a simple algorithm for accumulating overlapping Gaussians while keeping alpha blending between surfaces. Experimental setup shows benefits of the proposed approach in terms of rendering quality.
The quality of the camera calibration is of major importance for evaluating progresses in novel view synthesis, as a 1-pixel error on the calibration has a significant impact on the reconstruction quality. While there is no ground truth for real scenes, the quality of the calibration is assessed by the quality of the novel view synthesis. This paper proposes to use a 3DGS model to fine tune calibration by backpropagation of novel view color loss with respect to the cameras parameters. The new calibration alone brings an average improvement of 0.4 dB PSNR on the dataset used as reference by 3DGS. The fine tuning may be long and its suitability depends on the criticity of training time, but for calibration of reference scenes, such as Mip-NeRF 360, the stake of novel view quality is the most important.
3D Gaussian Splatting (3DGS) proposes an efficient solution for novel view synthesis. Its framework provides fast and high-fidelity rendering. Although less complex than other solutions such as Neural Radiance Fields (NeRF), there are still some challenges building smaller models without sacrificing quality. In this study, we perform a careful analysis of 3DGS training process and propose a new optimization methodology. Our Better Optimized Gaussian Splatting (BOGausS) solution is able to generate models up to ten times lighter than the original 3DGS with no quality degradation, thus significantly boosting the performance of Gaussian Splatting compared to the state of the art.
Light Field Image (LFI) has garnered remarkable interest and fascination due to its burgeoning significance in immersive applications. Although the abundant information in LFIs enables a more immersive experience, it also poses a greater challenge for Light Field Image Quality Assessment (LFIQA), especially when reference information is inaccessible. In this paper, inspired by the holistic visual perception of high-dimensional LFIs and neuroscience studies on the Human Visual System (HVS), we propose a novel Blind Light Field image quality assessment metric by exploring MultiPlane Texture and Multilevel Wavelet Information, abbreviated as MPT-MWI-BLiF. Specifically, considering the texture sensitivity of the secondary visual cortex (V2), we first convert LFIs into multiple individual planes and capture textural variations from these planes. Then, the statistical histogram of textural variations for all planes is calculated as holistic textural variation features. In addition, motivated by the fact that neuronal responses in the visual cortex are frequency-dependent, we simulate this visual perception process by decomposing LFIs into multilevel wavelet subbands with Four-Dimensional Discrete Haar Wavelet Transform (4D-DHWT). After that, the subband geometric features of first-level 4D-DHWT subbands and the coefficient intensity features of second-level 4D-DHWT subbands are computed respectively. Finally, we combine all the extracted quality-aware features and employ the widely-used Support Vector Regression (SVR) to predict the perceptual quality of LFIs. To fully validate the effectiveness of the proposed metric, we perform extensive experiments on five representative LFIQA databases with two cross-validation methods. Experimental results demonstrate the superiority of the proposed metric in quality evaluation, as well as its low time complexity compared to other state-of-the-art metrics. The full code will be publicly available at https://github.com/ZhengyuZhang96/MPT-MWI-BLiF
Light Field Image (LFI) records both angular and spatial information and provides immersive experiences for observers by rendering a scene from multiple perspectives. To cope with the resolution limitations of capture hardware, LFI angular reconstruction and spatial super-resolution are two widely-used methods, but they can also induce some special types of distortions, especially when two methods are adopted in combination. To this end, new challenges have been brought in assessing the quality of these distorted LFIs. In this paper, firstly, we conduct subjective experiments to evaluate the distorted LFI quality and present a novel perceptual quality assessment database with the associated subjective quality scores. Specifically, the proposed database focuses on the distortions introduced by deep learning-based LFI angular reconstruction and spatial super-resolution methods, individually and multiplely. Besides, in the case of multiple distortions, the adoption order of two distortions is taken into consideration. Further, our database presents three types of LFIs that suffer from distortions: real-world, dense synthesis, and sparse synthesis. As a result, the quality of distorted LFIs was subjectively assessed by 32 valid observers using the Pairwise Comparison (PC) protocol. Secondly, we develop a novel objective No-Reference (NR) metric for LFI quality evaluation, based on the features extracted from spatial gradients, angular-spatial statistics, and binocular disparity. Finally, a benchmark of the proposed metric and numerous state-of-the-art quality assessment metrics on the proposed database is presented. Experimental results demonstrate the superiority of the proposed metric over most existing metrics in various aspects. The proposed database and metric will be publicly available at https://github.com/ZhengyuZhang96/IETR-LFI.
HTTP adaptive streaming (HAS) has emerged as a prevalent approach for over-the-top (OTT) video streaming services due to its ability to deliver a seamless user experience. A fundamental component of HAS is the bitrate ladder, which comprises a set of encoding parameters (e.g., bitrate-resolution pairs) used to encode the source video into multiple representations. This adaptive bitrate ladder enables the client’s video player to dynamically adjust the quality of the video stream in real-time based on fluctuations in network conditions, ensuring uninterrupted playback by selecting the most suitable representation for the available bandwidth. The most straightforward approach involves using a fixed bitrate ladder for all videos, consisting of pre-determined bitrate-resolution pairs known as one-size-fits-all . Conversely, the most reliable technique relies on intensively encoding all resolutions over a wide range of bitrates to build the convex hull , thereby optimizing the bitrate ladder by selecting the representations from the convex hull for each specific video. Several techniques have been proposed to predict content-based ladders without performing a costly, exhaustive search encoding. This article provides a comprehensive review of various convex hull prediction methods, including both conventional and learning-based approaches. Furthermore, we conduct a benchmark study of several handcrafted- and deep learning (DL)-based approaches for predicting content-optimized convex hulls across multiple codec settings. The considered methods are evaluated on our proposed large-scale dataset, which includes 300 UHD video shots encoded with software and hardware encoders using three state-of-the-art video standards, including AVC/H.264, HEVC/H.265, and VVC/H.266, at various bitrate points. Our analysis provides valuable insights and establishes baseline performance for future research in this field ( Dataset URL : https://nasext-vaader.insa-rennes.fr/ietr-vaader/datasets/br_ladder ).
In the present work, an end-to-end approach is proposed for recovering an RGB-D scene representation directly from a hologram using its phase space representation. The proposed method involves four steps. First, a set of silhouette images is extracted from the hologram phase space representation. Second, a minimal 3D volume that describes these silhouettes is extracted. Third, the extracted 3D volume is decomposed into horizontal slices, and each slice is processed using a neural network to generate a coarse estimation of the scene geometry. Finally, a third neural network is employed to refine the estimation for higher precision applications. Experimental results demonstrate that the proposed approach yields faster and more accurate results compared to numerical reconstruction-based methods. Moreover, the obtained RGB-D representation can be directly utilized for alternative applications such as motion estimation.
Neural fields, also known as implicit neural representations (INRs), have shown a remarkable capability of representing, generating, and manipulating various data types, allowing for continuous data reconstruction at a low memory footprint. Though promising, INRs applied to video compression still need to improve their rate-distortion performance by a large margin, and require a huge number of parameters and long training iterations to capture high-frequency details, limiting their wider applicability. Resolving this problem remains a quite challenging task, which would make INRs more accessible in compression tasks. We take a step towards resolving these shortcomings by introducing neural representations for videos NeRV++, an enhanced implicit neural video representation, as more straightforward yet effective enhancement over the original NeRV decoder architecture, featuring separable conv2d residual blocks (SCRBs) that sandwiches the upsampling block (UB), and a bilinear interpolation skip layer for improved feature representation. NeRV++ allows videos to be directly represented as a function approximated by a neural network, and significantly enhance the representation capacity beyond current INR-based video codecs. We evaluate our method on UVG, MCL JVC, and Bunny datasets, achieving competitive results for video compression with INRs. This achievement narrows the gap to autoencoder-based video coding, marking a significant stride in INR-based video compression research.
By recording scenes from multiple viewpoints, Light Field Image (LFI) encompasses both angular and spatial information, thereby offering users a more immersive experience. Since LFIs may be distorted at various stages from acquisition to visualization, Light Field Image Quality Assessment (LFIQA) is of vitally important to monitor the potential impairments of LFI quality. However, existing objective LFIQA metrics fail to establish a reasonable correlation between spatial and angular information in LFIs, especially ignoring the imbalance problem of large spatial variations and subtle angular variations, which results in unsatisfactory quality evaluation performance. To alleviate this imbalance, in this paper, we propose a novel Blind LFIQA metric based on Angular-Spatial Effect Modeling, abbreviated as ASEM-BLiF. Specifically, the proposed metric consists of two branches. In the principal branch, we first present an Angular Effect Modeling (AEM) module to capture the angular information independently of spatial information. Based on AEM, we further design an Angular-Spatial Quality Learning (ASQL) module to model the local angular-spatial effect and establish the global relationship between different local regions for quality assessment via Transformer. In the auxiliary branch, a Discriminative Region Selection (DRS) module is proposed for auxiliary learning to improve the learning efficiency and prediction accuracy from a local perspective. Moreover, we present a Dynamic Weighting Loss (DWLoss) to achieve an optimal balance between principal and auxiliary learning throughout training. To demonstrate the effectiveness of the proposed metric, extensive experiments are conducted on five publicly available LFIQA databases with a variety of metrics. The experimental results show that compared to our previous work DeeBLiF, the current state-of-the-art LFIQA metric, our proposed ASEM-BLiF metric achieves 5.67%, 7.75%, 5.96%, 4.44%, and 0.33% SROCC performance improvements in quality assessment on the Win5-LID, NBU-LF1.0, LFDD, VALID10bit, and SHU databases, respectively. The code will be publicly available.
Accurately estimating 3D optical flow in computer-generated holography poses a challenge due to the scrambling of 3D scene information during hologram acquisition. Therefore, to estimate the scene motion between consecutive frames, the scene geometry should be recovered first. Recent studies have demonstrated that a 3D RGB-D representation can be extracted from an input hologram with relatively low error under well-chosen numerical reconstruction parameters. However, limited attention has been given to how the produced error can impact the flow estimation algorithms. Therefore, in this study, we evaluate different learning/non-learning methodologies for recovering 3D scene geometry. Next, we analyze the types of distortions produced by these methods and attempt to minimize estimation error using spatial and temporal constraints. Finally, we compare the performance of several state-of-the-art methods to estimate the 3D optical flow vectors on the recovered sequence of RGB-D images.
Abstract Video Coding for Machines (VCM) is gaining momentum in applications like autonomous driving, industry manufacturing, and surveillance, where the robustness of machine learning algorithms against coding artifacts is one of the key success factors. This work complements the MPEG/JVET standardization efforts in improving the resilience of deep neural network (DNN)-based machine models against such coding artifacts by proposing the following three advanced fine-tuning procedures for their training: (1) the progressive increase of the distortion strength as the training proceeds; (2) the incorporation of a regularization term in the original loss function to minimize the distance between predictions on compressed and original content; and (3) a joint training procedure that combines the proposed two approaches. These proposals were evaluated against a conventional fine-tuning anchor on two different machine tasks and datasets: image classification on ImageNet and semantic segmentation on Cityscapes. Our joint training procedure is shown to reduce the training time in both cases and still obtain a 2.4% coding gain in image classification and 7.4% in semantic segmentation, whereas a slight increase in training time can bring up to 9.4% better coding efficiency for the segmentation. All these coding gains are obtained without any additional inference or encoding time. As these advanced fine-tuning procedures are standard-compliant, they offer the potential to have a significant impact on visual coding for machine applications.
Despite the growing interest for Holography, there is a lack of publicly available three-dimensional hologram sequences for the evaluation of video codecs with inter-frame compression mechanisms such as motion estimation and compensation. In this paper, we report the first large-scale dataset containing 18 holographic videos computed with three different resolutions and pixel pitches. By providing the color and depth images corresponding to each hologram frame, our dataset can be used in additional applications such as the validation of 3D scene geometry retrieval or deep learning-based hologram synthesis methods. Altogether, our dataset comprises 5400 pairs of RGB-D images and holograms, totaling more than 550 GB of data.
Motivated by the efficiency investigation of the Tranformer-based transform coding framework, namely SwinT-ChARM, we propose to enhance the latter, as first, with a more straightforward yet effective Tranformer-based channel-wise auto-regressive prior model, resulting in an absolute image compression transformer (ICT). Current methods that still rely on ConvNet-based entropy coding are limited in long-range modeling dependencies due to their local connectivity and an increasing number of architectural biases and priors. On the contrary, the proposed ICT can capture both global and local contexts from the latent representations and better parameterize the distribution of the quantized latents. Further, we leverage a learnable scaling module with a sandwich ConvNeXt-based pre/post-processor to accurately extract more compact latent representation while reconstructing higher-quality images. Extensive experimental results on benchmark datasets showed that the proposed adaptive image compression transformer (AICT) framework significantly improves the trade-off between coding efficiency and decoder complexity over the versatile video coding (VVC) reference encoder (VTM-18.0) and the neural codec SwinT-ChARM.
In recent years, neural image compression has garnered considerable attention from both research and industry. It has shown great promise in surpassing traditional methods in terms of rate-distortion performance through the development of end-to-end deep neural codecs. Despite these advancements, there is still room for improvement, particularly in reducing the coding rate while maintaining high reconstruction fidelity, especially in non-homogeneous textured image areas. Current models, including attention-based transform coding, also tend to have a higher number of parameters and longer decoding times. To address these challenges, we propose ConvNeXt-ChARM, an efficient ConvNeXt-based transform coding framework. It is coupled with a compute-efficient channel-wise auto-regressive prior that captures both global and local contexts from the hyper and quantized latent representations. Our architecture can be optimized end-to-end, fully leveraging context information to extract compact latent representations and achieve higher-quality image reconstructions. Experimental results conducted on four widely-used datasets demonstrate the effectiveness of ConvNeXt-ChARM. It consistently delivers significant BD-rate (PSNR) reductions, averaging 5.24% over the VVC reference encoder (VTM-18.0) and 1.22% over the state-of-the-art learned image compression method SwinT-ChARM. Additionally, we conduct model scaling studies to verify the computational efficiency of our approach. Furthermore, we perform objective and subjective analyses to highlight the performance gap between ConvNeXt, the next-generation ConvNet, and the Swin Transformer. Overall, our proposed ConvNeXt-ChARM framework showcases improved compression efficiency and reconstruction quality, establishing itself as a promising solution in the field of neural image compression.
The increasing demand for immersive experience has greatly promoted the quality assessment research of Light Field Image (LFI). In this paper, we propose an efficient deep discrepancy measuring framework for full-reference light field image quality assessment. The main idea of the proposed framework is to efficiently evaluate the quality degradation of distorted LFIs by measuring the discrepancy between reference and distorted LFI patches. Firstly, a patch generation module is proposed to extract spatio-angular patches and sub-aperture patches from LFIs, which greatly reduces the computational cost. Then, we design a hierarchical discrepancy network based on convolutional neural networks to extract the hierarchical discrepancy features between reference and distorted spatio-angular patches. Besides, the local discrepancy features between reference and distorted sub-aperture patches are extracted as complementary features. After that, the angular-dominant hierarchical discrepancy features and the spatial-dominant local discrepancy features are combined to evaluate the patch quality. Finally, the quality of all patches is pooled to obtain the overall quality of distorted LFIs. To the best of our knowledge, the proposed framework is the first patch-based full-reference light field image quality assessment metric based on deep-learning technology. Experimental results on four representative LFI datasets show that our proposed framework achieves superior performance as well as lower computational complexity compared to other state-of-the-art metrics.
Recently, HTTP adaptive streaming (HAS) has become a standard approach for over-the-top (OTT)-based video streaming services due to its ability to provide smooth streaming. In HAS, stream representations are encoded to target a specific bitrate providing a wide range of operating bitrates known as the bitrate ladder. In the past, a fixed bitrate ladder approach for all videos has been widely used. However, such a method does not consider video content, which can vary considerably in motion, texture, and scene complexity. Moreover, building a per-title bitrate ladder based on an exhaustive encoding is quite expensive due to the large encoding parameter space. Thus, alternative solutions allowing accurate and efficient per-title bitrate ladder prediction are in great demand. On the other hand, self-attention-based architectures have achieved tremendous performance in large language models (LLMs) and particularly vision transformers (ViTs) in computer vision tasks. Therefore, this paper investigates ViT’s capabilities in building an efficient bitrate ladder without performing any encoding process. We provide the first in-depth analysis of the prediction accuracy and the complexity overhead induced by the ViTs model in predicting the bitrate ladder on a large and diverse video dataset. The source code of the proposed solution and the dataset will be made publicly available.
Light Field Image Quality Assessment (LF-IQA) is vitally important to facilitate the development of immersive technologies. However, current state-of-the-art LF-IQA metrics still struggle to handle Light Field Image (LFI) with massive data in an efficient manner. To cope with this challenge, we propose a simple yet effective Blind LF-IQA metric based on Spatio-Angular Textural Variation, named SATV-BLiF. Given a distorted LFI, we first apply Local Binary Pattern (LBP) operator to measure the textural variation in the spatial and angular domains respectively. Then the generated spatial and angular textural matrices are merged and further transformed into statistical textural histogram features. Finally, Support Vector Regression (SVR) is employed to construct a nonlinear mapping function between the statistical textural histogram features and the perceptual quality score of the distorted LFI. Experimental results on three representative light field databases show that the proposed metric achieves state-of-the-art quality evaluation performance, while having much lower complexity than the existing No-Reference (NR) LF-IQA metrics. The code of the proposed SATV-BLiF metric is available at https://github.com/ZhengyuZhang96/SATV-BLiF.
In this work, we introduce a novel approach for depth estimation in a computer-generated hologram by employing horizontal segmentation of the reconstruction volume instead of conventional vertical segmentation. The reconstruction volume is divided into horizontal slices and each slice is processed using a residual U-net architecture to identify in-focus lines, enabling determination of the slice's intersection with the 3D scene. The individual slice results are then combined to generate a dense depth map of the scene. Our experiments demonstrate the effectiveness of our method, with improved accuracy, faster processing times, lower graphics processing unit (GPU) utilization, and smoother predicted depth maps than existing state-of-the-art models.
Going beyond traditional 2D imaging is not only an emerging trend of imaging technology, but also the key to a more immersive user experience. Light Field Image (LFI) is a typical high-dimensional imaging format, and the quality evaluation of which is very challenging but necessary. In this article, we propose a novel Pseudo Video-based Blind quality assessment metric for Light Field image (PVBLiF). In contrast to most previous Light Field Image Quality Assessment (LF-IQA) metrics, in which different types of 2D representations derived from LFI are used for quality assessment indirectly, our metric exploits a more intuitive 3D representation, named Pseudo Video Block Sequence (PVBS), to evaluate the perceptual quality of LFI. For this purpose, we first divide the LFI into a massive number of non-overlapping PVBSs, which simultaneously contain spatial and angular information of LFI. Then, we propose a novel network (named PVBSNet) based on Convolutional Neural Networks (CNNs) to extract the spatio-angular features of PVBS and further evaluate the PVBS quality. The proposed PVBSNet consists of four stages: multi-information division, intra-feature extraction, cross-feature fusion, and quality regression. Finally, a Saliency- and Variance-guided Pooling (SVPooling) method is presented to integrate all the PVBS quality into the overall quality of LFI. The proposed PVBLiF metric has been extensively evaluated on three widely-used LFI datasets: Win5-LID, NBU-LF1.0, and SHU. Experimental results demonstrate that our proposed PVBLiF metric outperforms state-of-the-art metrics and is capable of highly approximating the performance of human observers.