Achieving personalized Quality of Experience (QoE) prediction and optimization is a long-term goal in media quality assessment, and recent research has shifted from predicting Mean Opinion Scores (MOS) toward modeling and predicting the quality judgments of individual subjects. In this work, we propose a model of individual image quality assessment that explicitly accounts for two key components of human visual perception: the spatial allocation of attention and the extraction of perceptual information from different regions of the visual scene. The model describes individual quality perception as a multi-stage process in which perceptual information and computational attention are iteratively combined and updated, followed by a probabilistic decision stage. From a theoretical perspective, the proposed formulation is defined at an abstract level and does not rely on any specific neural network architecture, and it provides a framework for explaining how computational attention and perceptual information interact to produce individual quality judgments. To illustrate the practical applicability of the proposed model, we instantiate it using a Vision Transformer (ViT) architecture as one possible implementation, training one AI-based Observer (AIO) per subject with observer-specific learned parameters. Experimental results show that the resulting ViT-based AIOs (i) compare favorably with state-of-the-art deep CNN-based individual models; (ii) outperform state-of-the-art no-reference image quality assessment metrics when used to predict individual opinion scores; (iii) tend to allocate greater attention to regions with severe quality degradation.
A recent research direction is focused on training Deep Neural Networks (DNNs) to replicate individual subject assessments of media quality. These DNNs are referred to as Artificial Intelligence-based Observers (AIOs). An AIO is designed to simulate, in real-time, the quality ratings of a specific individual, enabling an automatic quality assessment that accounts for subjects characteristics and preferences. Training AIOs is a promising but challenging research area due to the greater noise in individual raw opinion scores compared to the Mean Opinion Score. Effective learning from noisy labels necessitates the training of complex models on large-scale datasets. Unfortunately, this is challenging for AIOs as the media quality assessment community lacks extensive datasets that include individual opinion scores. To address the complexity of the task, we first created a dataset comprising two million samples, with synthetic labels derived from human annotation. We then trained a customized network for image quality assessment, named Multi-Distortion ResNet50 (MDResNet50), on this dataset. The weights of the MDResNet50 were subsequently utilized to initialize the learning process of each AIO, thereby avoiding the need to train a complex model from scratch on a small-scale dataset with raw individual opinion scores. Computational experiments show that our approach significantly advances the state-of-the-art in the AIO research. In particular: (i) we demonstrate through a simulation the ability of AIOs to mimic two well-known behavioral characteristics of a subject, i.e., bias and inconsistency, when scoring the media quality; (ii) we train and release DNN-based AIOs that, compared to the state-of-the-art, exhibit a higher performance with a statistical significance in assessing multiple image distortions; (iii) we train AIOs that more accurately mimic the sensitivity of real subjects to noise and color saturation and also better predict the opinion score distribution compared to the state-of-the-art AIOs.
Several approaches have been proposed to estimate quality in subjective experiments while highlighting peculiar subject behaviors. However, there is some room for improvement in existing approaches, both in terms of robustness to noise and the ability to accurately indicate several peculiar subject behaviors in subjective experiments. This work advances the state-of-the-art in three main directions: i) A new approach to estimate the subjective quality from noisy ratings is proposed and is shown to be more robust to noise than are four state-of-the-art approaches; ii) a novel subject scoring model is proposed that makes it possible to highlight several peculiar behaviors typically observed in subjective experiments; and iii) our proposed probabilistic subject scoring model results from the proof of a theorem, whereas in previous approaches a probabilistic scoring model is assumed a priori . This represents an important first step toward models supported by a stronger theoretical foundation. Numerical experiments conducted on several datasets highlight the effectiveness of our proposal.
Predicting the quality perception of an individual subject instead of the mean opinion score is a new and very promising research direction. Deep Neural Networks (DNNs) are suitable for such prediction but the training process is particularly data demanding due to the noisy nature of individual opinion scores. We propose a human-in-the-loop training process using multiple cycles of a human voting, DNN training, and inference procedure. Thus, opinion scores on individualized sets of images were progressively collected from each observer to refine the performance of their DNN. The results of computational experiments demonstrate the effectiveness of our approach. For future research and benchmarking, five DNNs trained to mimic five observers are released together with a dataset containing the 1500 opinion scores progressively gathered from each of these observers during our training cycles.
Unlike traditional objective approaches aimed at MOS prediction, subjective experiments provide individual opinion scores that allow, for instance, to estimate the distribution of users’ opinion scores. Unfortunately, the current literature is lacking objective quality assessment approaches that simulate the process of a subjective test. Therefore, this work focuses on modeling an individual subject through a deep CNN that, once trained, is expected to mimic the subject in terms of quality perception; for this reason, we call it “Artificial Intelligence-based Observer” (AIO). Several AIOs, modeling subjects with different characteristics, can be derived and used to simulate the process of a subjective test, thus yielding a more complete objective quality assessment. However, the training of the AIOs is hindered by two major issues: (i) the lack of training sets containing a large number of individual opinion scores; (ii) the noisy nature of individual opinion scores used as ground truth. To overcome these issues, we motivate a two-step learning approach. During the first learning step, the architecture of the well-known ResNet50 is appropriately modified and its initial weights are updated using a large scale synthetically annotated dataset of JPEG compressed images created for quality assessment purpose. This yields a new deep CNN called JPEGResNet50 that can accurately evaluate the quality of JPEG compressed images. The second learning step, conducted on a subjectively annotated dataset, refines the generic perceptual quality features already learned by the JPEGResNet50 to derive the AIO of each subject. Extensive computational experiments show the potential and effectiveness of our approach.
Despite several approaches to recover the ground truth subjective quality score from noisy individual ratings in subjective experiments have been explored in the literature, there is still room for improvement, in particular in terms of robustness to noise. This paper proposes a new approach that combines the traditional maximum likelihood estimation framework with a newly proposed regularization term, based on information theory concepts, that is meant to underweight surprising ratings of the quality of a given stimulus, looked at as a noise manifestation, in the final analytical expression of the recovered subjective quality. Computational experiments show the higher robustness to noise of our proposal when compared to three state-of-the-art methods.
The media quality assessment research community has traditionally been focusing on developing objective algorithms to predict the result of a typical subjective experiment in terms of Mean Opinion Score (MOS) value. However, the MOS, being a single value, is insufficient to model the complexity and diversity of human opinions encountered in an actual subjective experiment. In this work we propose a complementary approach for objective media quality assessment that attempts to more closely model what happens in a subjective experiment in terms of single observers and, at the same time, we perform a qualitative analysis of the proposed approach while highlighting its suitability. More precisely, we propose to model, using neural networks (NNs) , the way single observers perceive media quality. Once trained, these NNs, one for each observer, are expected to mimic the corresponding observer in terms of quality perception. Then, similarly to a subjective experiment, such NNs can be used to simulate the users’ single opinions, which can be later aggregated by means of different statistical indicators such as average, standard deviation, quantiles, etc. Unlike previous approaches that consider subjective experiments as a black box providing reliable ground truth data for training, the proposed approach is able to consider human factors by analyzing and weighting individual observers. Such a model may therefore implicitly account for users’ expectations and tendencies, that have been shown in many studies to significantly correlate with visual quality perception. Furthermore, our proposal also introduces and investigates an index measuring how much inconsistency there would be if an observer was asked to rate many times the same stimulus. Simulation experiments conducted on several datasets demonstrate that the proposed approach can be effectively implemented in practice and thus yielding a more complete objective assessment of end users’ quality of experience.
From the beginnings of ITU-T H.261 to H.265 (HEVC), each new video coding standard has aimed at halving the bitrate at the same perceptual quality by redundancy and irrelevancy reduction. Each improvement has been explained by comparably small changes in the video coding toolset. This contribu-tion aims at starting the Quality of Experience (QoE) analysis of the accumulated improvements over the last thirty years. Based on an overview of the changes in the coding tools, we analyze the changes in the quan-tized residual information. Visual comparison and sta-tistical measures are performed and some interpreta-tions are provided towards explaining how irrelevancy reduction may have led to such a huge reduction in bitrate. The interpretation of the results in terms of QoE paves the way towards an understanding of the coding tools in terms of visual quality. It may help in understanding how the irrelevancy reduction has been improved over the decades. Understanding how the differences of the residuals relate to known or yet un-known properties of the human visual system, may en-able a closer collaboration between perception research and video compression research.
Several video quality metrics (VQMs) have been proposed in many publications to predict how humans perceive video quality. It is common to observe significant disagreements amongst the quality predictions of these VQMs for the same video sequence. Following an extensive literature search, we found no publicised work that has investigated if such disagreements convey useful information on the accuracy of VQMs. Herein, a measure for quantifying the disagreement between VQMs is proposed. A small-scale subjective study is carried out to assess the effectiveness of our proposal. In particular, the proposed disagreement measure is shown to be extremely effective in determining whether the quality of any given processed video sequence (PVS) can be accurately predicted by the VQMs. This type of information is particularly useful for identifying video sequences that are likely to degrade the end-user’s quality of experience (QoE). Our proposal is also useful in selecting the most effective PVSs to be employed in a subjective test. We show that the proposed disagreement measure can be effectively predicted from bitstream features. This establishes a link between the capability to accurately assess the quality of a PVS and the way it is encoded. In addition, an analysis is conducted to compare the performances of some well-known and widely used open-source metrics and two proprietary metrics. The two proprietary metrics are used by a large media company for enhancing its delivery pipeline. The outcome of this comparison highlights the suitability of the open-source VQM, Video Multi-method Assessment Fusion (VMAF), as a good benchmark quality measure for both the industrial and academic environments.
There is a continuing demand for objective measures that predict perceived media quality. Researchers are developing new methods for mapping technical parameters of digital media to the perceived quality. It is quite common to use machine learning algorithms for these purposes especially deep learning algorithms, which need large amounts of data for training. In this paper, we aim towards getting more training data with recent types of distortions. Instead of doing expensive subjective experiments, we evaluate the reuse of previously published, well-known image datasets with subjective annotation. In this contribution, the procedure of mapping Mean Opinion Scores (MOS) from an already published subjectively annotated dataset with older codecs to new codecs is presented. In particular, we map from Joint Photographic Experts Group (JPEG) distortions to newer High Efficiency Video Coding (HEVC) distortions. We have used values of three different objective methods as a connection between these two different distortion types. In order to investigate the significance of our approach, subjective verification tests were designed and conducted. The design goals led to two types of experiments, i.e. Pair Comparison (PC) test and Absolute Category Rating (ACR) test, in which 40 participants provided their opinion. Results of the subjective experiments indicate that it may be possible to use information gained from older datasets to describe the perceived quality of more recent compression algorithms.
Subjective experiments are important for developing objective Video Quality Measures (VQMs). However, they are time-consuming and resource-demanding. In this context, being able to reuse existing subjective data on previous video coding standards to train models capable of predicting the perceptual quality of video content processed with newer codecs acquires significant importance. This paper investigates the possibility of generating an HEVC encoded Processed Video Sequence (PVS) in such a way that its perceptual quality is as similar as possible to that of an AVC encoded PVS whose quality has already been assessed by human subjects. In this way, the perceptual quality of the newly generated HEVC encoded PVS may be annotated approximately with the Mean Opinion Score (MOS) of the related AVC encoded PVS. To show the effectiveness of our approach, we compared the performance of a simple and low complexity but yet effective no reference hybrid model trained on the data generated with our approach with the same model trained on data collected in the context of a pristine subjective experiment. In addition, we merged seven subjective experiments such that they can be used as one aligned dataset containing either original HEVC bitstreams or the newly generated data explained in our proposed approach. The merging process accounts for the differences in terms of quality scale, chosen assessment method and context influence factors. This yields a large annotated dataset of HEVC sequences that is made publicly available for the design and training of no reference hybrid VQMs for HEVC encoded content.
Subjective experiments are considered the most reliable way to assess the perceived visual quality. However, observers’ opinions are characterized by large diversity: in fact, even the same observer is often not able to exactly repeat his first opinion when rating again a given stimulus. This makes the Mean Opinion Score (MOS) alone, in many cases, not sufficient to get accurate information about the perceived visual quality. To this aim, it is important to have a measure characterizing to what extent the observed or predicted MOS value is reliable and stable. For instance, the Standard deviation of the Opinions of the Subjects (SOS) could be considered as a measure of reliability when evaluating the quality subjectively. However, we are not aware of the existence of models or algorithms that allow to objectively predict how much diversity would be observed in subjects’ opinions in terms of SOS. In this work we observe, on the basis of a statistical analysis made on several subjective experiments, that the disagreement between the quality as measured by means of different objective video quality metrics (VQMs) can provide information on the diversity of the observers’ ratings on a given processed video sequence (PVS). In light of this observation we: i) propose and validate a model for the SOS observed in a subjective experiment; ii) design and train Neural Networks (NNs) that predict the average diversity that would be observed among the subjects’ ratings for a PVS starting from a set of VQMs values computed on such a PVS; iii) give insights into how the same NN based approach can be used to identify potential anomalies in the data collected in subjective experiments.
The last decades witnessed an increasing number of works aiming at proposing objective measures for media quality assessment, i.e. determining an estimation of the mean opinion score (MOS) of human observers. In this contribution, we investigate a possibility of modeling and predicting single observer’s opinion scores rather than the MOS. More precisely, we attempt to approximate the choice of one single observer by designing a neural network (NN) that is expected to mimic that observer behavior in terms of visual quality perception. Once such NNs (one for each observer) are trained they can be looked at as “virtual observers” as they take as an input information about a sequence and they output the score that the related observer would have given after watching that sequence. This new approach allows to automatically get different opinions regarding the perceived visual quality of a sequence whose quality is under investigation and thus estimate not only the MOS but also a number of other statistical indexes such as, for instance, the standard deviation of the opinions. Large numerical experiments are performed to provide further insight into a suitability of the approach.
The coding tools used in image and video encoders aim at high perceptual quality for low bitrates. Analyzing the results of the encoders in terms of quantization parameter, image partitioning, prediction modes or residuals may provide important insight into the link between those tools and the human perception. As a first step, this contribution analyzes the possibility to transcode reference images of three well-known image databases, i.e. IRCCyN/IVC, LIVE and TID2013, from their original, older formats to HEVC; thus creating a homogeneous database of 327 HEVC encoded images accompanied with bitstream parameters and values obtained from objective and subjective assessments. Secondly, it analyzes some of the HEVC intra coding parameters regarding their influence on the image quality by using machine learning, namely Support Vector Machine - Regression.
Typically, the measurement of the Quality of Experience for video sequences aims at a single value, in most cases the Mean Opinion Score (MOS). Predicting this value using various algorithms has been widely studied. However, deviation from the MOS is often handled as an unpredictable error. The approach in this contribution estimates intervals of video quality instead of the single valued MOS. Well-known video quality estimators are fused together to output a lower and upper border for the expected video quality, on the basis of a model derived from a well-known subjectively annotated dataset. Results on different datasets provide insight on the suitability of the well-known estimators for this particular approach.
In this study, we evaluate an interaction sequence performed by six modalities consisting of desktop-based (DB) and virtual reality (VR) environments using different input devices. For the given study, we implemented a vertical prototype of a first person shooter (FPS) game scenario, focusing on the genre-defining point-and-shoot mechanic. We introduce measures to evaluate the success of the according interaction sequence (times for target acquisition, pointing, shooting, overall net time, and number of shots) and conduct experiments to record and compare the users' performances. We show that interacting using head-tracking for landscape-rotation is performing similarly to the input of a screen-centered mouse and also yielded shortest times in target acquisition and pointing. Although using head-tracking for target acquisition and pointing was most efficient, subjects rated the modality using head-tracking for target acquisition and a 3DOF Controller for pointing best. Eye-tracking (ET) yields promising results, but calibration issues need to be resolved to enhance reliability and overall user experience.
Subjective quality assessment is a necessary activity to validate objective measures or to assess the performance of innovative video processing technologies. However, designing and performing comprehensive tests requires expertise and a large effort especially for the execution part. In this work we propose a methodology that, given a set of processed video sequences prepared by video quality experts, attempts to reduce the number of subjective tests by selecting a subset with minimum size which is expected to yield the same conclusions of the larger set. To this aim, we combine information coming from different types of objective quality metrics with clustering and machine learning algorithms that perform the actual selection, therefore reducing the required subjective assessment effort while trying to preserve the variety of content and conditions needed to ensure the validity of the conclusions. Experiments are conducted on one of the largest publicly available subjectively annotated video sequence dataset. As performance criterion, we chose the validation criteria for video quality measurement algorithms established by the International Telecommunication Union.
The training and performance analysis of objective video quality assessment algorithms is complex due to the huge variety of possible content classes and transmission distortions. Several secondary issues such as free parameters in machine learning algorithms and alignment of subjective datasets put an additional burden on the developer. In this paper, three subsequent steps are presented to address such issues. First, the content and coding parameter space of a large-scale database is used to select dedicated subsets for training objective algorithms. This aims at providing a method for selecting the most significant contents and coding parameters from all imaginable combinations. In the practical case where only a limited set is available, it also helps us to avoid redundancy in the training subset selection. The second step is a discussion on performance measures for algorithms that employ machine-learning methods. The particularity of the performance measures is that the quality of the training and verification datasets is taken into consideration. Common issues that often use existing measures are presented, and improved or complementary methods are proposed. The measures are applied to two examples of no-reference objective assessment algorithms using the aforementioned subsets of the large-scale database. While limited in terms of practical applications, this sandbox approach of objectively predicting an objectively evaluated video sequences allows for eliminating additional influence factors from subjective studies. In the third step, the proposed performance measures are applied to the practical case of training and analyzing assessment algorithms on readily available subjectively annotated image datasets. The presentation method in this part of the paper can also be used as an exemplified recommendation for reporting in-depth information on the performance. Using this presentation method, future publications presenting newly developed quality assessment algorithms may be significantly improved.
This work presents a framework to facilitate reproducibility of research in video quality evaluation. Its initial version is built around the JEG-Hybrid database of HEVC coded video sequences. The framework is modular, organized in the form of pipelined activities, which range from the tools needed to generate the whole database from reference signals up to the analysis of the video quality measures already present in the database. Researchers can re-run, modify and extend any module, starting from any point in the pipeline, while always achieving perfect reproducibility of the results. The modularity of the structure allows to work on subsets of the database since for some analysis this might be too computationally intensive. To this purpose, the framework also includes a software module to compute interesting subsets, in terms of coding conditions, of the whole database. An example shows how the framework can be used to investigate how the small differences in the definition of the widespread PSNR metric can yield very different results, discussed in more details in our accompanying research paper Aldahdooh et al. (0000). This further underlines the importance of reproducibility to allow comparing different research work with high confidence. To the best of our knowledge, this framework is the first attempt to bring exact reproducibility end-to-end in the context of video quality evaluation research. (C) 2017 The Authors. Published by Elsevier B.V.
Visual discomfort is an important factor that influences viewing experience in immersive multimedia, for example, 3DTV and VR. With the added value of depth, the novel perceptual experience, visual discomfort is not an easy task for observers to evaluate. In this study, we investigate how the subjective methodology affects the test results in 3DTV condition. Two subjective visual discomfort experiments were conducted. One used the Pair Comparison (PC) method and the other used the Absolute-Category Rating (ACR) method. The results demonstrated that PC method had more powerful discriminability. For a difficult perceptualrelated tasks, such as visual discomfort in our study, PC was more easy to understand and conduct for the observers which led to reliable results. It also showed some very important but usually ignored conclusions on the subjective experiment, i.e., for measuring the perceived visual discomfort, the observer's judgment behavior might be affected by the test methodology.