Video compression is a fundamental topic in the visual intelligence, bridging visual signal sensing/capturing and high-level visual analytics. The broad success of artificial intelligence (AI) technology has enriched the horizon of video compression into novel paradigms by leveraging end-to-end optimized neural models. In this survey, we first provide a comprehensive and systematic overview of recent literature on end-to-end optimized learned video coding, covering the spectrum of pioneering efforts in both uni-directional and bi-directional prediction based compression model designation. We further delve into the optimization techniques employed in learned video compression (LVC), emphasizing their technical innovations, advantages. Some standardization progress is also reported. Furthermore, we investigate the system design and hardware implementation challenges of the LVC inclusively. Finally, we present the extensive simulation results to demonstrate the superior compression performance of LVC models, addressing the question that why learned codecs and AI-based video technology would have with broad impact on future visual intelligence research.
Traditional in the wild image quality assessment (IQA) models are generally trained with the quality labels of mean opinion score (MOS), while missing the rich subjective quality information contained in the quality ratings, for example, the standard deviation of opinion scores (SOS) or even distribution of opinion scores (DOS). In this paper, we propose a novel IQA method named RichIQA to explore the rich subjective rating information beyond MOS to predict image quality in the wild. RichIQA is characterized by two key novel designs: (1) a three-stage image quality prediction network which exploits the powerful feature representation capability of the Convolutional vision Transformer (CvT) and mimics the short-term and long-term memory mechanisms of human brain; (2) a multi-label training strategy in which rich subjective quality information like MOS, SOS and DOS are concurrently used to train the quality prediction network. Powered by these two novel designs, RichIQA is able to predict the image quality in terms of a distribution, from which the mean image quality can be subsequently obtained. Extensive experimental results verify that the three-stage network is tailored to predict rich quality information, while the multi-label training strategy can fully exploit the potentials within subjective quality rating and enhance the prediction performance and generalizability of the network. RichIQA outperforms state-of-the-art competitors on multiple large-scale in the wild IQA databases with rich subjective rating labels. The code of RichIQA will be made publicly available on GitHub.
The rapid advancement of artificial intelligence (AI) technology has led to the prioritization of standardizing the processing, coding, and transmission of video using neural networks. To address this priority area, the Moving Picture, Audio, and Data Coding by Artificial Intelligence (MPAI) group is developing a suite of standards called MPAI-EEV for "end-to-end optimized neural video coding." The aim of this AI-based video standard project is to compress the number of bits required to represent high-fidelity video data by utilizing data-trained neural coding technologies. This approach is not constrained by how data coding has traditionally been applied in the context of a hybrid framework. This paper presents an overview of recent and ongoing standardization efforts in this area and highlights the key technologies and design philosophy of EEV. It also provides a comparison and report on some primary efforts such as the coding efficiency of the reference model. Additionally, it discusses emerging activities such as learned Unmanned-Aerial-Vehicles (UAVs) video coding which are currently planned, under development, or in the exploration phase. With a focus on UAV video signals, this paper addresses the current status of these preliminary efforts. It also indicates development timelines, summarizes the main technical details, and provides pointers to further points of reference. The exploration experiment shows that the EEV model performs better than the state-of-the-art video coding standard H.266/VVC in terms of perceptual evaluation metric.
During the past decade, the Unmanned-Aerial-Vehicles (UAVs) have attracted increasing attention due to their flexible, extensive, and dynamic space-sensing capabilities. The volume of video captured by UAVs is exponentially growing along with the increased bitrate generated by the advancement of the sensors mounted on UAVs, bringing new challenges for on-device UAV storage and air-ground data transmission. Most existing video compression schemes were designed for natural scenes without consideration of specific texture and view characteristics of UAV videos. In this work, we first contribute a detailed analysis of the current state of the field of UAV video coding. Then we propose to establish a novel task for learned UAV video coding and construct a comprehensive and systematic benchmark for such a task, present a thorough review of high quality UAV video datasets and benchmarks, and contribute extensive rate-distortion efficiency comparison of learned and conventional codecs after. Finally, we discuss the challenges of encoding UAV videos. It is expected that the benchmark will accelerate the research and development in video coding on drone platforms.
Considering multidimensional structure of the multi-exposure images, a new Tensor product and Tensor-singular value decomposition based Multi-Exposure image Fusion (TT-MEF) method is proposed. The main innovation of this work is to explore a new feature representation of multi-exposure images in the new tensor domain and design the fusion strategy on this basis. Specifically, the luminance and the chrominance channels are fused separately to maintain color consistency. For the luminance fusion, the luminance channel of multi-exposure images is divided into two parts, that is, de-mean term and mean term. The de-mean term is represented as a tensor to extract the feature. Then, the tensor product and tensor-singular value decomposition (T-SVD) are used to design a tensor feature extractor. Furthermore, a fusion strategy of the de-mean term is presented according to the visual saliency model, and a fusion strategy of the mean term is defined by the local and the global visual weights to control counterpoise between the local and global luminance. For the chrominance fusion, a new fusion strategy is also designed by the tensor product and T-SVD, similar to the luminance fusion. Finally, the fused image is obtained by combining the luminance and chrominance fusion. Experimental results show that the proposed TT-MEF method generally outperforms the existing state-of-the-art in terms of subjective visual quality and objective evaluation.
Sparse representation has been shown to be highly correlated with the visual perception of natural images, which can be characterized by a linear combination of neuronal responses in the visual cortex. Divisive normalization transform (DNT) has been proven to be an effective method in reducing statistical and perceptual dependencies for nonlinear properties in primary visual cortex. In this paper, we develop a divisively normalized sparse coding scheme, aiming to further bridge the gap between sparse representation and human visual perception. We show that such a scheme is perceptually meaningful for representing visual signals, with which the pixel-domain image representation and processing tasks can be feasibly and efficiently achieved in the divisively normalized sparse-domain. Specifically, we develop a sparse-domain similarity (SDS) index for perceptual quality evaluation, where the DNT is employed for transforming image signals into a perceptually uniform space. Furthermore, the proposed SDS index is employed to optimize the sparse coding process when representing natural images. The experimental results indicate that the SDS can provide accurate and consistent predictions of perceived image quality, and the performance of sparse coding can be significantly improved in terms of both objective and subjective quality evaluations.
Screen content images (SCIs) have been rapidly and widely applied in interactive multimedia applications. The problem of quality assessment for SCIs is an interesting research topic. Most of the existing methods use subjective and independent features in gray domain to predict the image quality, which cannot comprehensively characterize the image properties or lack unified mathematical explanation for SCIs. To address these problems, we propose a novel blind quality assessment method based on macro-micro modeling of tensor domain dictionary for SCIs in this article. In the proposed method, the tensor decomposition is explored first to avoid the loss of color information, and then a target dictionary is learned more effectively with the principal components. Furthermore, a macro-micro model is established to characterize the micro and macro features in the target dictionary space, which can provide a systematic mathematical interpretation for feature extraction. For the micro features, a log-normal pooling scheme is designed to enhance the effectiveness of feature aggregation by analyzing the particularity of the statistical distribution of sparse codes. Additionally, the statistical properties are mainly discussed and studied based on the Bernoulli law of large numbers, and then a reliable macro feature is generated to describe the relationship between the statistical distribution and quality degradation of SCIs. Experimental results determined by using three public SCI databases show that the proposed method can perform better than relevant existing methods in the prediction of the visual quality of SCIs, especially in terms of the generalization for distortion type and interpretability for feature generation.
Localizing an autonomous vehicle in real-time is critical for robust autonomous driving. As a standard approach, the map-based localization is robust and fast; however, it is expensive to create and maintain a large-scale high-definition map. In this paper, we propose an online localization technique based on the vehicle-to-vehicle communication and traffic landmark detection; called collaborative localization. This can potentially serve as a new complement to the standard localization solutions. We theoretically show that multiple vehicles with multiple traffic landmarks would significantly improve the localization performance. We then propose a practical algorithm, which leverages graph matching to handle practical issues, such as traffic landmark association. The experimental results validate the potential of the proposed methods.
Multimedia hardware still cannot accommodate the demand for large amounts of visual data. Without the generation of high-quality video bitstreams, limited hardware capabilities will continue to stifle the advancement of multimedia technologies. Thorough grounding in coding is needed so that applications such as MPEG-4 and JPEG 2000 may come to fruition. Image and Video Compression for Multimedia Engineering provides a solid, comprehensive understanding of the fundamentals and algorithms that lead to the creation of new methods for generating high quality video bit streams. The authors present a number of relevant advances along with international standards. New to the Second Edition A chapter describing the recently developed video coding standard, MPEG-Part 10 Advances Video Coding also known as H.264 Fundamental concepts and algorithms of JPEG2000 Color systems of digital video Up-to-date video coding standards and profiles Visual data, image, and video coding will continue to enable the creation of advanced hardware, suitable to the demands of new applications. Covering both image and video compression, this book yields a unique, self-contained reference for practitioners tobuild a basis for future study, research, and development.
Yun Q Shi合作论文数Department of Electrical and Computer Engineering;New Jersey Institute of Technology (NJIT)12
Kadir A. Peker合作论文数Research Laboratories8