The history of image quality has undergone a series of ages, each with distinct applications and levels of quantification, modelling, statistical analysis. This talk will cover perceptual quality for color images, video, and three-dimensional signals. Full-text article not available; see video presentation
Recent advancements in text-to-video and video-to-audio generative AI (GenAI) have spurred excitement around their ability to potentiate human creativity. However, the simulated imagery and audio produced by state-of-the-art GenAI models may contain artifacts and anomalies that diminish their realism. Our multimodal experiment investigates 1) the relative weighting of auditory versus visual information humans utilize to distinguish real user-generated content (UGC) from AI-generated content (AIGC) and 2) which visual artifacts produced by GenAI are most conspicuous to human observers. Observers completed a 2-interval forced-choice (2IFC) task with eye tracking, choosing which of two sequentially presented video clips was more realistic. Observers were more accurate at discerning real from fake imagery than real from fake audio. Eye movement analysis revealed that top-down motion-related artifacts, such as violations of biological motion, were more conspicuous than bottom-up artifacts, such as the intermingling of real-world textures with data compression artifacts.
Natural image statistics are well known to have a spatial frequency power spectra that has a 1/f^a behavior, with a typically stated as between 2 and 4. This indicates an invariance to scale. Further work has theorized how the visual system is tuned for such statistics in visual cortex (V1) [1]. Color image statistics also show an invariance to scale [2]. The luminance histogram is typically understood to be log normal with respect to luminance, although for HDR images, a subcomponent with skew toward much higher luminances is observed. Color statistics were initially described at the simplest level via the gray world hypothesis [3], but more details are now available, even at the hyperspectral [4]. The a power function for HDR was found to increase from the lower values of 2 to more typical values of 4 and 5 [5]. For temporal statistics, the data tends to be measured primarily for media, with a 1/f^a for scene cut statistics [6], and temporal frequency and temporal frequency for media with a focus on the motion statistics via optical flow [7]. Statistics for purely natural as well as human made environments (e.g., buildings and the resulting perspective geometry) have been studied, each having different orientation statistics [8]. The use of image statistics for standardized assessment of television power consumption was used to replace test targets, which were often detected and used to lower TV power consumption in well known cheating schemes. To prevent this, a short test video that had luminance statistics matching 48 hours of broadcast content was generated and used for TV power testing [9]. The highly adaptative nature of current TVs (power limiting, dual modulation, dynamic response) has motivated researchers to incorporate complex noise fields following natural image statistics into measurement targets [10,11]. One particular natural image statistic-based still image test target (dead leaves) is widely used in camera optics and sensor development. Algorithm development and testing for image and video processing has almost always been ad hoc, with a mixture of geometric test targets and hand selected test images, sometimes aiming to be corner cases, sometimes not. More recently, large data sets of images have been used to train various neural network models for tasks such as super resolution, bit rate compression, and dynamic range mapping. However, images are not ergodic, and possibly not even wide-sense stationary. We propose the use of imagery based on noise following the natural image statistics for spatio-chromatic (and temporal) to compactly probe the wide variety of image possibilities for algorithmic development, in addition to the existing uses for image capture and display analysis. While we don’t suggest replacing actual practical imagery, we believe such noise fields can augment image algorithm analysis. To address the problem of non-ergodicity, we allow the basic power term a in the natural image statistic model to vary over a large range in a video, such that it includes the extremes of white noise and low frequency gradients. We use color image statistic models that include decorrelated colors to generate the RGB video. We will present results for traditional adaptive data compression (with chromatic subsampling), as well as a more contemporary neural network approach (Neural Fields [12]) as applied to upscaling and denoising. We analyze the results both visually and through several recent color image quality models. Field DJ. Relations between the statistics of natural images and the response properties of cortical cells. J. Opt. Soc. Am. A, 1987; 4:2379-2394 C. Parraga, T. Troscianko, and D.J. Tolhurst (2002) spatiochromatic properties of natural images and human vision. Current Biology V 12 R. M. Evans, Method for correcting photographic color prints, US Patent 2,571,697 (1951) A. Chakrabarti and T. Zickler (2011) Statistics of real-world hyperspectral images CVPR R. Dror, A. Willsky, and E. Adelson (2004) statistical characterization of real-world illumination. JOV V4 J. Cutting (2019) Sequences in popular cinema generate inconsistent event segmentation. Attn. Percept. And Psycho. V 81. D. Lee, H. Ko, J. Kim, and A. Bovik (2021) On the space-time statistics of motion pictures. JOSA A V 38 #7 A. Torralba and A. Oliva (2003) Statistics of natural image categories, Network: Computational Neural Systems 14 391-412 International Electrotechnical Commission, IEC 62087:2008(E), “Methods of measurement for the power consumption of audio, video, and related Equipment. Kunkel T, Daly S. 57-1: Spatiotemporal Noise Targets Inspired by Natural Imagery Statistics. SID Symposium Digest of Technical Papers, 2020, 51:842-845. Kunkel, T, Friedrich, F. Utilizing advanced spatio-temporal backgrounds with dynamic test signals for high dynamic range display metrology. J Soc Inf Display. 2022; 30( 5): 423– 432. https://doi.org/10.1002/jsid.1125 Yiheng Xie1, Towaki Takikawa, Shunsuke Saito, Or Litany, Shiqin Yan, Numair Khan, Federico Tombari, James Tompkin, Vincent Sitzmann, Srinath Sridhar1, "Neural Fields in Visual Computing and Beyond", Eurographics / CGF State-of-the-Art Report, 2022.
The paper describes a design of a subjective experiment for testing the video quality of High Dynamic Range, Wide Color gamut (HDR-WCG) content at 4K resolution. Due to Covid, testing could not use a lab, so an at-home test procedure was developed. To aim for calibration despite not fully controlling the conditions and settings, we limited subjects to those who had a specific TV model, which we had previously calibrated in our labs. Moreover, we performed the experiment in the Dolby Vision mode (where the various enhancements of the TV are turned OFF by default). A browser approach was used which took control of the TV, and ensure the content was viewed at the native resolution of the TV (e.g., dot-on-dot mode). In addition, we know that video imagery is not ergodic, and there is wide variability in types of low levels features (sharpness, noise, motion, color volume, etc.) that affect both TV and visual system performance. So, a large number of test clips was used (30) and the content was specifically chosen to stress key features. The obtained data is qualitatively similar to an in-lab study and is subsequently used to evaluate several existing objective quality metrics.
Efficiently transmitting media while maintaining or improving the quality of experience (QoE) poses a constant challenge. On the media playback side, there exists a variety of factors affecting the resulting QoE perceived by end users, such as device characteristics, viewing conditions, and individual visual acuity. Understanding the characteristics of playback-side context can provide opportunities for optimizing media delivery tailored to each user. — We present a novel approach to estimating viewers’ visible frequency ranges by creating hybrid images in a quantitatively controlled manner. The proposed approach, which can be applied to adaptive bitrate (ABR) streaming optimization, combines low- and high-frequency components of two different images providing multiscale image perception. The proposed method is designed to adjust the hybrid image size and the cutoff frequencies of lowpass and highpass filters of hybrid image synthesis according to the given frequency range under test. This method can be used to estimate the minimum video resolution for maintaining maximum QoE, in any viewing conditions and devices. — Our experimental results reveal a strong correlation between the minimum resolution obtained from hybrid image perception tasks and the traditional just noticeable difference (JND) of real images of different resolutions when presented side-by-side. These JND tasks exhibit a wider spread both in the measured minimum resolution and elapsed time to complete the tasks, potentially due to the complexity of the cognitive process and the variability caused by contents and subjects. In contrast, the hybrid image technique achieves more stable results with much less effort from subjects.
SMPTE was formed in 1916 as narrative performance was ushered onto the technological stage allowing for mass distribution and transmission. While there is no need to list those advantages, there is one weakness with audiovisual media as compared to the traditional stage, which media still has not overcome. This is real-time audience feedback, and the ability to adjust the performance based on differing audience reactions. This paper motivates the use of biosensors in media by highlighting the problem of signal loss due to playback technology. A metadata system is proposed that allows creatives to steer signal modifications as a function of audience emotion and cognition as determined by biosensor analysis This is needed because today's audiovisual ecosystem includes such a wide variety of playback devices that the audience's experience can differ substantially for the same source content. Metadata for narrative and emotional expectation as inserted by creatives during the content production stages combines with the assessment to adapt the rendering. As a result, the system allows for creative intent to be scalable as best as possible across many types of playback systems in a manner analogous to real-time stage presentation.
The motivation for use of biosensors in audiovisual media is made by highlighting problem of signal loss due to wide variability in playback devices. A metadata system that allows creatives to steer signal modifications as a function of audience emotion and cognition as determined by biosensor analysis.
Modern content is rapidly shifting from the standard dynamic range (SDR) color volume defined in ITU-R BT.709 towards the high dynamic range and wide color gamut (HDR-WCG) color volume defined in ITU-R BT.2100. For optimizing delivery channels of this content it is crucial to have high performance video quality metrics. However, existing subjective databases used to design and qualify such video quality metrics do not provide adequate test coverage for professional HDR-WCG video content. In this paper we describe a subjective video quality study aimed to be relevant for existing as well as near future HDR-WCG. — Due to the Covid pandemic, the experiment was designed for remote testing, rather than a typical laboratory environment. We carefully analyzed the subjective results and concluded that it is possible to obtain comparable reliability of data in our remote testing environment with a controlled laboratory environment. This opens the door for large-scale video quality testing, enabling continual refinement of video quality metrics, reducing the turn-around time for subjective testing, and eliminating bias in the training data by accessing a much larger and more diverse set of users. — We applied the subjective data set to compare existing image/video quality metrics such as VMAF, HDR-VDP and HDR-VQM. Furthermore, we identified several ways the VMAF metric can be adjusted to improve its accuracy on HDR-WCG content.
High dynamic range (HDR) and wide color gamut imagery has an established video ecosystem, spanning image capture to encoding and display. This drives the need for evaluating how image quality is affected by the multitudes of ecosystem parameters. The simplest quality metrics evaluate color differences on a pixel-by-pixel basis. In this article, we evaluate a series of these color difference metrics on four HDR and three standard dynamic range publicly available distortion databases consisting of natural images and subjective scores. We compare the performance of the well-established CIE L*a*b* metrics (Delta E-00, Delta E-94) alongside two HDR-specific metrics (Delta E-Z [J(z)a(z)b(z)], Delta E-ITP [ICTCP]) and a spatial CIE L*a*b* extension (Delta E00S). We also present a novel spatial extension to Delta E-ITP derived by optimizing the opponent color contrast sensitivity functions. We observe that this advanced metric, Delta EITPSC, outperforms the other color difference metrics, and we quantify the improved performance with the steps of metric advancement.
There is a long tradition of using simple geometric synthetic test patterns for display and video metrology, and a somewhat newer tradition of using pictorial imagery for overall quality assessment, such as for subjective studies and demonstrations. Each of these key approaches have disadvantages. The synthetic targets can be stress testers that have no bearing on practical usage of the imaging system and act to drive up costs, while pictorial imagery can easily fail on robustness assessment and lack a foundation for analysis. To alleviate the shortcomings of both approaches, we propose a system of noise field videos designed using specific statistical parameters as a foundation for video test targets for display assessment, and also provide a tool to generate such videos.
This paper presents an approach to predicting image quality by spatially filtering images before generating color difference maps with pixel‐based color difference metrics. The resulting difference maps can then be pooled across the whole image. This approach was originally developed for CIELAB color space under the name S‐CIELAB. We extend this approach to use the recently developed ICTCP color space to improve the prediction accuracy for high dynamic range and wide color gamut images. The filtering is based on the chromatic and achromatic contrast sensitivity function of the human visual system. Our results on four existing subjective image quality databases containing high dynamic range and wide color gamut images show substantial improvements at low computational cost, outperforming existing color difference metrics.
There are an increasing number of databases describing subjective quality responses for HDR (high dynamic range) imagery with various distortions. The dominant distortions across the databases are those that arise from video compression, which are primarily perceived as achromatic, but there are some chromatic distortions due to 422 and other chromatic sub-sampling. Tone mapping from the source HDR levels to various levels of reduced capability SDR (standard dynamic range) are also included in these databases. While most of these distortions are achromatic, tone-mapping can cause changes in saturation and hue angle when saturated colors are in the upper hull of the of the color space. In addition, there is one database that specifically looked at color distortions in an HDR-WCG (wide color gamut) space. From these databases we can test the improvements to well-known quality metrics if they are applied in the newly developed color perceptual spaces (i.e., representations) specifically designed for HDR and WCG. We present results from testing these subjective quality databases to computed quality using the new color spaces of Jzazbz and ICTCP, as well as the commonly used SDR color space of CIELAB.
We improve High Dynamic Range (HDR) Image Quality Assessment (IQA) using a full reference approach that combines results from various quality metrics (HDR-CQM). We combine metrics designed for different applications such as HDR, SDR and color difference measures in a single unifying framework and non-linearly combine the scores from different quality metrics using support vector machine regression. To improve performance and reduce complexity, we use the Sequential Forward Selection technique to select a subset of metrics from a list of quality metrics. We evaluate the performance on two publicly available databases with different types of distortion and demonstrate improved performance using HDR-CQM as compared to several existing IQA metrics. We also show the generality and robustness of our approach using cross-database evaluation.
High dynamic range (HDR) and wide color gamut (WCG) imagery is now mainstream across content creation. Thus, having a reliable way of evaluating HDR systems is essential. One common way of measuring the quality of an HDR system is measuring the color errors introduced along the imaging pipeline. Unfortunately, development of color difference metrics has been based on test patches, as opposed to natural imagery, with evaluation mostly limited to standard dynamic range (SDR) databases. In this paper, we evaluate several color difference metrics on five publicly available HDR distortion databases consisting of natural images and subjective scores. Distortion types include lower frequency distortions such as those from tone mapping as well as higher-frequency distortions resulting from compression artifacts by the various compression schemes such as Joint Photographic Experts Group (JPEG). We analyze these databases with the well-established CIE $\text{L}^\star \text{a}^\star \text{b}^\star $ metrics ( $\Delta E_{00}, \Delta E_{94}$ ) as well as two HDR-specific metrics: $\Delta E_{z} (j_{z}a_{z}b_{z})$ and $\Delta E_{ITP} (IC_{T}C_{P})$ . To quantify the performance, we use four standard performance evaluation procedures. We observe that $\Delta E_{ITP}$ outperforms the other color difference metrics on four out of five databases. We also perform statistical analysis, which demonstrates the effectiveness of using $\Delta E_{ITP}$ to assess the overall image quality of HDR/WCG content. Weaknesses of all metrics for one database suggest that more advanced spatio-chromatic visual models are worth pursuing.
We present an approach to predict overall high dynamic range (HDR) display quality as a function of key HDR display parameters. Subjective experiments on a high quality HDR display tested five key HDR display parameters: (maximum luminance, minimum luminance, color gamut, bit-depth and local contrast). Two models were explored- a physical model solely based on physically measured display characteristics and a perceptual model that transforms physical parameters using human vision system (HVS) models. For the perceptual model, we used a family of metrics based on a recently published color space (ICTCP), as well as an estimate of the display point spread function. To predict the overall visual quality ratings, we compared linear regression and various machine learning (ML) techniques. We find that the perceptual model is better at predicting subjective quality than the physical model and that ML techniques are better at prediction than linear regression. We also investigate the significance and contribution of each display parameter for the combined model. Further, we explore how well these models perform when applied to display capabilities outside of the training data set, both in terms of extrapolation and interpolation. The use of the perceptual transforms particularly helps with extrapolation, and without their tempering effects, the ML-based models can produce wildly unrealistic quality predictions.
We present a framework for assessing the interaction of display resolution, display size, and video compression in the context of visual acuity. We review retinal topology in the context of several kinds of acuity: reading (optotype) acuity; grating acuity; and hyperacuity. We show that display resolution and field‐of‐view combinations correspond to each level of visual acuity. We also show how quantization of spatial‐frequency coefficients in video compression relates to each level of visual acuity.
The perceived discrepancy between continuous motion as seen in nature and frame-by-frame exhibition on a display, sometimes termed judder, is an integral part of video presentation. Over time, content creators have developed a set of rules and guidelines for maintaining a desirable cinematic look under the restrictions placed by display technology without incurring prohibitive judder. With the advent of novel displays capable of high brightness, contrast, and frame rates, these guidelines are no longer sufficient to present audiences with a uniform viewing experience. In this work, we analyze the main factors for perceptual motion artifacts in digital presentation and gather psychophysical data to generate a model of judder perception. Our model enables applications like matching perceived motion artifacts to a traditionally desirable level and maintain a cinematic motion look.
When calibrating wide‐gamut displays, a colorimetric match may result in perceived color error due to inaccuracies of the CIE 1931 observer. In this study, methods of correcting values measured on wide gamut displays to perceptually equivalent values on standard gamut displays were analyzed using visual color matching data.
Andreas Schilling合作论文数Wilhelm-Schickard-Institut für Informatik, Fachbereich Informatik, Mathematisch-Naturwissenschaftliche Fakultät, Universität Tübingen2