Considerable progress has been made toward developing still picture perceptual quality analyzers that do not require any reference picture and that are not trained on human opinion scores of distorted images. However, there do not yet exist any such completely blind video quality assessment (VQA) models. Here, we attempt to bridge this gap by developing a new VQA model called the video intrinsic integrity and distortion evaluation oracle (VIIDEO). The new model does not require the use of any additional information other than the video being quality evaluated. VIIDEO embodies models of intrinsic statistical regularities that are observed in natural vidoes, which are used to quantify disturbances introduced due to distortions. An algorithm derived from the VIIDEO model is thereby able to predict the quality of distorted videos without any external knowledge about the pristine source, anticipated distortions, or human judgments of video quality. Even with such a paucity of information, we are able to show that the VIIDEO algorithm performs much better than the legacy full reference quality measure MSE on the LIVE VQA database and delivers performance comparable with a leading human judgment trained blind VQA model. We believe that the VIIDEO algorithm is a significant step toward making real-time monitoring of completely blind video quality possible.
Successful video quality analysers make use of a reference video to compare against or by training on a database of human rated distorted videoes as priors, both of which are either not available or difficult to obtain in many practical scenarios. Although efforts have been made towards designing still picture quality analyzers that are `completely blind' and do not require any prior training on, or exposure to, distorted images or human opinions of them [1], we are attempting to fill an important but challenging gap by designing a `completely blind' video naturalness analyser. The principle of this new approach is based on the regularties observed in time-frequency relationships of natural vidoes across time. Our experimental results on the LIVE VQA (video quality assessment) database [2] show that, even with no prior knowledge, the new VQA algorithm performs better than the full reference (FR) quality measure PSNR. The approach is very lean in computational expense which makes it a very good candidate for real time signal processing applications.
We propose a no reference (NR) video quality assessment (VQA) model. Recently, ‘completely blind’ still picture quality analyzers have been proposed that do not require any prior training on, or exposure to, distorted images or human opinions of them. We have been trying to bridge an important but difficult gap by creating a ‘completely blind’ VQA model. The principle of this new approach is founded on intrinsic statistical regularities that are observed in natural vidoes. This results in a video ‘quality analyzer’ that can predict the quality of distorted videos without any external knowledge about the pristine source, anticipated distortions or human judgments. Hence, the model is zero shot. Experimental results show that, even with such paucity of information, the new VQA algorithm performs better than the full reference (FR) quality measure PSNR on the LIVE VQA database. It is also fast and efficient. We envision that the proposed method is an important step towards making real time monitoring of ‘completely blind’ video quality feasible.
We propose a family of image quality assessment (IQA) models based on natural scene statistics (NSS), that can predict the subjective quality of a distorted image without reference to a corresponding distortionless image, and without any training results on human opinion scores of distorted images. These `completely blind' models compete well with standard non-blind image quality indices in terms of subjective predictive performance when tested on the large publicly available `LIVE' Image Quality database.
We define the new idea of blind image repair as a process of correcting one or more different and unknown types of distortions afflicting an image. These distortions could introduce linear or non-linear degradations, compression artifacts, noise, etc., or combinations of these. Thus the concept encompasses denoising, deblurring, deblocking, deringing, and other post-acquisition image improvement processes that address distortions. The problem is distortion-blind when the natures of the distortion processes are unknown prior to analyzing the image. Towards solving this problem, we describe a new framework for repairing an image that has undergone an unknown set of distortions, based on identifying the distortion(s) present in the image (if any) and applying possibly multiple distortion-specific image repair algorithms. Our philosophy is based on the principle that the task of general purpose image repair is one of agglomeration, i.e., the algorithm should embody multiple high-performing distortion-specific repair modules such that seamless general purpose image repair is achieved. Our proposed framework – the GEneral-purpose No-reference Image Improver (GENII) – enables the design of algorithms that are blind to distortion type as well as to distortion parameters, and only requires as input the distorted image to be repaired. The GENII framework is modular and easily extensible to image repair problems beyond those considered here. GENII operates by using natural scene statistic models to identify distortion, to perceptually optimize the distortion parameter(s), to assess the quality of the intermediate repaired images, and to perceptually optimize the repair processes. We explain the general purpose image repair framework and one specific realization, dubbed GENII-1, which assumes that the image has been affected by one or more of four possible distortion types.The performance of GENII-1 is evaluated on 4000 distorted images, and shown to deliver substantial improvements in both quantitative and qualitative visual quality.
This work studies connections between image naturalness and perceived image quality. Specifically we define an Image Naturalness Index that quantifies intrinsic naturalness of images using features learned from a representative database of natural images. These features derive from models of early visual processing that lead to statistically regular processed images. We have observed that image distortions disrupt statistical image naturalness and that humans are highly sensitive to these disruptions in distorted images they observe. The naturalness index is derived by selecting patches from natural images, then collecting relevant NSS features from these patches to construct a natural image model. A multivariate Gaussian (MVG) distribution is used to characterize them. The 'naturalness' of an arbitrary test image, whose quality needs to be evaluated is expressed as the distance of the MVG model obtained from natural images to the MVG fit on the set of the same NSS features extracted from the patches of the test image. When applied to distorted images, the index is able to achieve image quality assessment (IQA) performance (in terms of correlation with human subjective judgments) comparable to leading full reference IQA algorithms such the Structural Similarity (SSIM) Index and much better than the Mean Squared Error (MSE) which demonstrates its relevance with respect to perceptual distortion sensitivity. This method is different from prior IQA approaches, all of which required both similar NSS models as well as exposure to both distorted images and human judgments of them, replying instead only on a simple natural scene statistic (NSS) model. Conversely, the Image Naturalness Index performs well when constructed using perceptually relevant NSS features extracted from a corpus of naturalistic, undistorted features, but does not function well when using other types of features (such as SIFT features or edges) or if the features are extracted from distorted images. Meeting abstract presented at VSS 2013
Stereoscopic/3D image and video quality assessment (IQA/VQA) has become increasing relevant in today's world, owing to the amount of attention that has recently been focused on 3D/stereoscopic cinema, television, gaming, and mobile video. Understanding the quality of experience of human viewers as they watch 3D videos is a complex and multi-disciplinary problem. Toward this end we offer a holistic assessment of the issues that are encountered, survey the progress that has been made towards addressing these issues, discuss ongoing efforts to resolve them, and point up the future challenges that need to be focused on. Important tools in the study of the quality of 3D visual signals are databases of 3D image and video sets, distorted versions of these signals and the results of large-scale studies of human opinions of their quality. We explain the construction of one such tool, the LIVE 3D IQA database, which is the first publicly available 3D IQA database that incorporates 'true' depth information along with stereoscopic pairs and human opinion scores. We describe the creation of the database and analyze the performance of a variety of 2D and 3D quality models using the new database. The database as well as the algorithms evaluated are available for researchers in the field to use in order to enable objective comparisons of future algorithms. Finally, we broadly summarize the field of 3D QA focusing on key unresolved problems including stereoscopic distortions, 3D masking, and algorithm development.
An important aim of research on the blind image quality assessment (IQA) problem is to devise perceptual models that can predict the quality of distorted images with as little prior knowledge of the images or their distortions as possible. Current state-of-the-art "general purpose" no reference (NR) IQA algorithms require knowledge about anticipated distortions in the form of training examples and corresponding human opinion scores. However we have recently derived a blind IQA model that only makes use of measurable deviations from statistical regularities observed in natural images, without training on human-rated distorted images, and, indeed without any exposure to distorted images. Thus, it is "completely blind." The new IQA model, which we call the Natural Image Quality Evaluator (NIQE) is based on the construction of a "quality aware" collection of statistical features based on a simple and successful space domain natural scene statistic (NSS) model. These features are derived from a corpus of natural, undistorted images. Experimental results show that the new index delivers performance comparable to top performing NR IQA models that require training on large databases of human opinions of distorted images. A software release is available at http://live.ece.utexas.edu/research/quality/niqe_release.zip.
Natural scene statistic (NSS) models are effective tools for formulating models of early visual processing. One area where NSS models have been successful is predicting human responses to image distortions, or image quality assessment (IQA) by quantifying unnaturalness introduced by distortions. Recent Blind IQA models use NSS features to form predictions of human judgments of distorted image quality without having available corresponding undistorted reference images. Successful learning blind models have previously been developed that learn to accurately predict human opinions of image quality by training them on databases of distorted images and associated human opinion scores. We introduce new NSS feature based blind IQA models that require even less information to attain good results. If human opinion scores of distorted images are not available, but a database of distorted images is, then opinion-less blind IQA models can be created that perform well. We have also found it possible to design blind IQA models without any source of prior information other than a database of distortionless "exemplar" images. An algorithm derived from such a completely blind model has only the distorted image to be quality-assessed available. Our new blind IQA models (Fig. 1) follow four processing steps (Fig. 2). Images are decomposed by an energy compacting filter bank then divisive normalized, yielding responses well-modeled as NSS. Either NSS features alone, or both NSS and distorted image statistic (DSS) features are used to create distributions of visual words. Quality prediction is expressed in terms of the Kullback-Leibler divergence between the distributions of visual words from distorted images and from the space of exemplar images. Both opinion blind and completely blind models compete well with standard non-blind metrics such as mean squared error (MSE) when tested on a large public IQA database (Tables 1 and 2). Meeting abstract presented at VSS 2012
Subjective studies have been conducted in the past to obtain human judgments of visual quality on distorted images in order, among other things, to benchmark objective image quality assessment (IQA) algorithms. Existing subjective studies primarily have records of human ratings on images that were corrupted by only one of many possible distortions. However, the majority of images that are available for consumption are corrupted by multiple distortions. Towards broadening the corpora of records of human responses to visual distortions, we recently conducted a study on two types of multiply distorted images to obtain human judgments of the visual quality of such images. Further, we compared the performance of several existing objective image quality measures on the new database and analyze the effects of multiple distortions on commonly used quality-determinant features and on human ratings.
A natural scene statistics (NSS) based blind image denoising approach is proposed, where denoising is performed without knowledge of the noise variance present in the image. We show how such a parameter estimation can be used to perform blind denoising by combining blind parameter estimation with a state-of-the-art denoising algorithm.1 Our experiments show that for all noise variances simulated on a varied image content, our approach is almost always statistically superior to the reference BM3D implementation in terms of perceived visual quality at the 95% confidence level.
Computing relative or absolute range (egocentric distance) is difficult because, of course, neither is specified in any direct way by the 2D retinal image. If, however, there was a relationship between range and luminance or color, perhaps it could be exploited to yield fast, initial estimates of range from the retinal image per se. We studied the statistical dependence between range (and disparity) contrast and luminance contrast across random point-pairs in natural scenes, and found that changes in range and luminance are highly dependent. We collected high resolution range maps of natural scenes co-registered with luminance (RGB) images using a Riegl terrestrial scanner, co-mounted camera, and in-house software. Various alternative preprocessing stages were used to simulate the early stages of visual processing (e.g. foveation). Our basic approach was to randomly sample pairs of points in the scenes to determine if the change in range or luminance or both exceeded some criterion. We then 1) compared the conditional density of range edges given luminance edges to the (unconditioned) density of range edges and 2) compared the joint distribution of range and luminance contrast to the product of their marginal distributions. We found a robust statistical dependence between range and luminance. Additionally, we computed difference surface maps (between the joint distributions and product-of-marginals predicted by independence). These difference surfaces reveal which regions of luminance and range change exhibit the strongest statistical dependencies. The statistical dependence between luminance and range allows the construction of models where one can assign a probability of occurrence of a range edge given a luminance edge at a particular point in a scene. In principle, such a mechanism could also be used by biological visual system to serve as priors when reconstructing the 3D environment from 2D image data. Meeting abstract presented at VSS 2012
We propose a natural scene statistic-based distortion-generic blind/no-reference (NR) image quality assessment (IQA) model that operates in the spatial domain. The new model, dubbed blind/referenceless image spatial quality evaluator (BRISQUE) does not compute distortion-specific features, such as ringing, blur, or blocking, but instead uses scene statistics of locally normalized luminance coefficients to quantify possible losses of “naturalness” in the image due to the presence of distortions, thereby leading to a holistic measure of quality. The underlying features used derive from the empirical distribution of locally normalized luminances and products of locally normalized luminances under a spatial natural scene statistic model. No transformation to another coordinate frame (DCT, wavelet, etc.) is required, distinguishing it from prior NR IQA approaches. Despite its simplicity, we are able to show that BRISQUE is statistically better than the full-reference peak signal-to-noise ratio and the structural similarity index, and is highly competitive with respect to all present-day distortion-generic NR IQA algorithms. BRISQUE has very low computational complexity, making it well suited for real time applications. BRISQUE features may be used for distortion-identification as well. To illustrate a new practical application of BRISQUE, we describe how a nonblind image denoising algorithm can be augmented with BRISQUE in order to perform blind image denoising. Results show that BRISQUE augmentation leads to performance improvements over state-of-the-art methods. A software release of BRISQUE is available online: http://live.ece.utexas.edu/research/quality/BRISQUE_release.zip for public use and evaluation.
We develop a robust framework for natural scene statistic (NSS) model based blind image quality assessment (IQA). The robustified IQA model utilizes a robust statistics approach based on L-moments. Such robust statistics based approaches are effective when natural or distorted images deviate from assumed statistical models, and achieves better prediction performance on distorted images relative to human subjective judgments. We also show how robustifying the model makes IQA approach resilient against deviation in model assumptions, small variations in the distortions and amount of data the model is trained on.
We performed a systematic evaluation of ‘visually lossless’ (VL) threshold selection for H.264/AVC (Advanced Video Coding) compressed natural videos spanning a wide range of content and motion. A psychovisual study was conducted using a two alternative forced choice task design, where by a series of reference vs. compressed video pairs were displayed to the subjects, where bit rates were varied to achieve a spread in the amount of compression. A statistical analysis was conducted on these data to estimate the VL threshold. Based on the visual thresholds estimated from the observed human ratings, we learn a mapping from ‘perceptually relevant’ statistical video features that capture visual lossless-ness, to statistically determined VL threshold. Using this VL threshold, we derive an H.264 compressibility index. This new Compressibility Index is shown to correlate well with human subjective judgments of VL thresholds. We have also made the code for compressibility index available online (Moorthy, A.K. and Bovik, A.C. (2010). H.264 Visually Lossless Compressibility Index (HVLCI), Software Release. http://live.ece.utexas.edu/research/quality/hvlci.zip.) for its use in practical applications and facilitate future research in this area.
We tracked the points-of-gaze of human observers as they viewed videos drawn from foreign films while engaged in two different tasks: (1) Quality Assessment and (2) Summarization. Each video was subjected to three possible distortion severities - no compression (pristine), low compression and high compression - using the H. 264 compression standard. We have analyzed these eye-movement locations in detail. We extracted local statistical features around points-of-gaze and used them to answer the following questions: (1) Are there statistical differences in variances of points-of-gaze across videos between the two tasks?, (2) Does the variance in eye movements indicate a change in viewing strategy with change in distortion severity? (3) Are statistics at points-of-gaze different from those at random locations? (4) How do local low-level statistics vary across tasks? (5) How do point-of-gaze statistics vary across distortion severities within each task?