With the widespread use of image processing technologies, objective image quality metrics are a fundamental and challenging problem. In this paper, we present a new No-Reference Image Quality Assessment (NR-IQA) algorithm based on visual attention modeling and a multivariate Gaussian distribution to predict the final quality score from the extracted features. Computational modeling of visual attention is performed to compute saliency maps at three resolution levels. At each level, distortions of the input image are extracted and weighted by the saliency maps in order to highlight degradations of visually attracting regions. The generated features are used by a probabilistic model to predict the final quality score. Experimental results demonstrate the effectiveness of the metric and show better performance when compared to well known NR-IQA algorithms.
With the rapid growth of multimedia applications and technologies, objective image quality assessment (IQA) became a topic of fundamental interest. No-Reference (NR) IQA algorithms are more suitable to real-world applications where the original image is not available. In order to be more consistent with human perception, this paper proposes a new NR-IQA metric where the input image is firstly decomposed to several frequency sub-bands which mimic the human visual system (HVS). Then, the statistical features are extracted from these frequency bands and used to fit a multivariate Gaussian distribution (MVGD). Finally, the model obtained by training predicts the quality of the input image. Experimental results demonstrate the method effectiveness and show its robustness when tested by different databases. Moreover, the predicted quality is more consistent with human perception.
Visual attention modeling is a very active research field and several image and video attention models have been proposed during the last decade. However, despite the conclusions drawn from various studies about the influence of human gazes by the presence of sound, most of the classical video attention models do not account for the multimodal nature of video (visual and auditory cues). In this paper, we propose an audiovisual saliency model with the aim to predict human gaze maps when exploring video content. The model, intended for videoconferencing, is based on the fusion of spatial, temporal and auditory attentional maps. Based on a real-time audiovisual speaker localization approach, the proposed auditory map is modulated depending of the nature of faces in the video, i.e. speaker or auditor. State-of-the-art performance measures have been used to compare the predicted saliency maps with the eye-tracking ground truth. The obtained results show the very good performance of the proposed model and a significant improvement compared to non-audio models.
published or not.The documents may come from teaching and research institutions in France or abroad, or from public or private research centers.L'archive ouverte pluridisciplinaire HAL, est destinée au dépôt et à la diffusion de documents scientifiques de niveau recherche, publiés ou non, émanant des établissements d'enseignement et de recherche français ou étrangers, des laboratoires publics ou privés. Dimension reduction-based attributes selection in no-reference learning-based image quality algorithms
With the rapid growth of image processing technologies, objective Image Quality Assessment (IQA) is a topic where considerable research effort has been made over the last two decades. IQA algorithms based on image structure have been shown to correlate well with Mean Opinion Scores (MOS). No-Reference (NR) image quality metrics are of fundamental interest as they can be embedded in practical applications This paper deals with a new NR-IQA metric based on natural scenes statistics. It proposes to model the best correlated statistics of seven well known no-reference image quality algorithms by a MultiVariate Gaus-sian Distribution (MVGD). A part of LIVE database is used with the associated DMOS to fit the MVGD model, namely Model Image Quality Index (MIQI). Hence, the quality of a distorted image is given by the DMOS that maximizes the multivariate Gaussian probability density function. Experimental results demonstrate the method effectiveness for a wide variety of distortions.
Visual attention modeling is a very active research area. During the last decade several image and video attention models have been proposed. Unfortunately, the majority of classical video attention models do not take into account the multimodal aspect of the video (visual and auditory cues). However, several studies have proven that human gazes are affected by the presence of the soundtrack. In this paper we propose an audiovisual saliency model that can predict the human gaze maps when exploring a conferencing or conversation videos. The model is based on the fusion of spatial, temporal and auditory attentional maps. Thanks to a real-time audiovisual speaker localization method, the proposed auditory maps are modulated by the enhanced saliency region of speakers compared to the other faces in the video. Classical visual attention measures have been used to compare the predicted saliency maps with the eye-tracking ground truth. Results of the proposed approach, using several fusion methods, show a good performance whatever the used spatial models.
No-reference image quality metrics are of fundamental interest as they can be embedded in practical applications. The main goal of this paper is to perform a comparative study of seven well known no-reference learning-based image quality algorithms. To test the performance of these algorithms, three public databases are used. As a first step, the trial algorithms are compared when no new learning is performed. The second step investigates how the training set influences the results. The Spearman Rank Ordered Correlation Coefficient (SROCC) is utilized to measure and compare the performance. In addition, an hypothesis test is conducted to evaluate the statistical significance of performance of each tested algorithm.
Objective and subjective quality assessment of images and videos have been an active research topic in the recent years. Multimedia technologies require new quality metrics and methodologies taking into account the fundamental differences in the human visual perception and the typical distortions of both video and audio modalities. Because of the increase of multimedia content platforms (Streaming, IPTV, OTT, etc) to delivering video content over the open Internet to a variety of devices; from TVs, to tablets, and smart phones, the Quality of Experience (QoE) may change. In this work we propose to evaluate the mutual interaction between audio and video and their influence on the perceived quality in case of streaming video applications. From one side, we carried out subjective experiments for assessing audio-only, video-only and audiovisual quality to create an audiovisual database that contains a wide range of degradations. From another side, a statistical analysis has been performed to investigate the influence of video resolution, viewing device and audio quality on perceived audiovisual quality. The Results show that audio quality plays a crucial role in the judgment of perceived quality especially when its quality is poor.
While objective and subjective quality assessment of images and video have been an active research topic in the recent years, multimedia technologies require new quality metrics and methodologies taking into account the fundamental differences in the human visual perception and the typical distortions of both video and audio modalities. Because of the importance of faces and especially the talking faces in the video sequences, this paper presents an audiovisual database that contains a different talking scenario. In addition to the video, the database also provides subjective quality scores obtained using a tailored single-stimulus test method (ACR). The resulting mean opinion scores (MOS) can be used to evaluate the performance of audiovisual quality metrics as well as for the comparison and for the design of new models.
Usual attention is an important mechanism of the human visual system. It allows reducing the amount of information to be processed and accelerates the overall process of vision. Several models for images and videos have been proposed in the literature with encouraging results. However, most existing saliency models do not take into account the multimodal aspect of the video (audio and image). In this paper, we propose to investigate the influence of audio on visual attention. From one side, we carried out eye-tracking experiments for recording subjects' eye movement when watching videos. From another side, eye positions (fixation duration) were used for computing and comparing video attention maps (through eye-movements) with and without audio using state-of-the-art measures. Results therefore showed that audio significantly affect the viewer's attention and consequently it must be taken into consideration for the development of any multimedia saliency model. The findings of these experiments have been used for the development of audiovisual saliency model based on talking face. Result obtained with our model were assessed using usual measurements and showed good performance with regards to ground truth.
No reference image quality metrics are of fundamental interest as they can be embedded in practical applications. This research domain is subject of intensive activities and numerous objective models have been proposed in literature. The main goal of this paper is to perform a comparative study of seven well known no-reference image quality algorithms. To test the performance of these algorithms, three public databases are used. The Spearman rank ordered correlation coefficient is utilized to measure and compare the performance. In addition, an hypothesis test is conducted to evaluate the statistical significance of performance of each tested algorithm.
Rate control algorithm adopted in H.264/AVC reference software shows several shortcomings that have been highlighted by different studies. For instance, in the baseline profile, the frame target bit-rate estimation assumes similar characteristics for all frames and the quantization parameter determination uses the Mean Absolute Difference for complexity estimation. Consequently, an inefficient bit allocation is performed leading to important quality variation of decoded sequences. A saliency-based rate-control is proposed in this paper to achieve bit-rate saving and improve perceived quality. The saliency map of each frame, simulating the human visual attention by a bottom-up approach, is used at the frame level to adjust the quantization parameter and at the macroblock level to guide the bit allocation process. Simulation results show that the proposed attentional model is well correlated to human behavior. When compared to JM15.0 reference software, at the frame level, the saliency map exploitation achieves bit-rate savings of up to 26%. At the MB level and under the same quality constraint, bit-rate improvement is up to 42% and buffer level variation is reduced by up to 71%.
Rate control is a critical issue in H.264/AVC video coding standard because it suffers from some shortcomings that make the bit allocation process not optimal. This leads to a video quality that may vary significantly from frame to frame. Our aim is to enhance the rate control efficiency in H.264/AVC baseline profile by handling two of its defects: the initial quantization parameter (QP) estimation for Intra-Frames (I-Frames) and the target number of bits determination for Inter-Frames (P-Frames) encoding. First, we propose a Rate-Quantization (R-Q) model for the I-Frame constructed empirically after extensive experiments. The optimal initial QP calculation is based on both target bit-rate and I-Frame complexity. The I-Frame target bit-rate is derived from the global target bit-rate by using a new non-linear model. Secondly, we propose an enhancement of the bit allocation process by exploiting frame complexity measures. The target number of bits determination for P-Frames is adjusted by combining two temporal measures: the first is a motion ratio based on actual bits used to encode previous frames; the second measure exploits the difference between two consecutive frames and the histogram of this difference. The simulation results, carried out using the JM15.0 reference software and the JVT-O016 rate control algorithm, show that the right choice of initial QP for I-Frame and first P-Frame allows improvement of both the bit-rate and peak signal-to-noise ratio (PSNR). Finally, the Inter-Frame bit allocation process further improves the bit-rates while keeping the same PSNR improvement (up to +1.33 dB/ +2 dB for QCIF/CIF resolutions). Moreover, this process reduces the buffer level variation leading to a more consistent quality of reconstructed videos. (c) 2012 SPIE and IS&T. [DOI: 10.1117/1.JEI.21.1.013013]
Rate control is a critical issue in H.264/AVC video coding standard because it suffers from some shortcomings that make the bit allocation process not optimal. This leads to a video quality that may vary significantly from frame to frame. Our aim is to enhance the rate control efficiency in H.264/AVC baseline profile by handling two of its defects: the initial quantization parameter (QP) estimation for Intra-Frames (I-Frames) and the target number of bits determination for Inter-Frames (P-Frames) encoding. First, we propose a Rate-Quantization (R-Q) model for the I-Frame constructed empirically after extensive experiments. The optimal initial QP calculation is based on both target bit-rate and I-Frame complexity. The I-Frame target bit-rate is derived from the global target bit-rate by using a new non-linear model. Secondly, we propose an enhancement of the bit allocation process by exploiting frame complexity measures. The target number of bits determination for P-Frames is adjusted by combining two temporal measures: the first is a motion ratio based on actual bits used to encode previous frames; the second measure exploits the difference between two consecutive frames and the histogram of this difference. The simulation results, carried out using the JM15.0 reference software and the JVT-O016 rate control algorithm, show that the right choice of initial QP for I-Frame and first P-Frame allows improvement of both the bit-rate and peak signal-to-noise ratio (PSNR). Finally, the Inter-Frame bit allocation process further improves the bit-rates while keeping the same PSNR improvement (up to +1.33 dB/+2 dB for QCIF/CIF resolutions). Moreover, this process reduces the buffer level variation leading to a more consistent quality of reconstructed videos.
Dans cet ouvrage les auteurs recensent les concepts fondamentaux et les dernieres avancees dans le domaine de l'acquisition, de la perception, du codage et du rendu des couleurs. Destine aux chercheurs et ingenieurs, aux etudiants en Master ou Doctorat, cet ouvrage dresse un etat de l'art sur les problematiques scientifiques et techniques soulevees par les differentes etapes de la chaine numerique couleur. Cet ouvrage aborde les aspects fondamentaux lies a la colorimetrie et a la physiologie, a la constance et a l'apparence des couleurs. Il traite aussi des aspects plus techniques lies aux capteurs et a la gestion des couleurs sur ecran. Une attention particuliere a ete egalement apportee a la notion de rendu des couleurs en synthese d'images. Au dela de la couleur, un etat de l'art approfondit est aussi mene sur le codage, la compression, la protection et la qualite d'images et de videos couleur.
Dans cet ouvrage les auteurs recensent les concepts fondamentaux et les dernieres avancees dans le domaine de l'acquisition, de la perception, du codage et du rendu des couleurs. Destine aux chercheurs et ingenieurs, aux etudiants en Master ou Doctorat, cet ouvrage dresse un etat de l'art sur les problematiques scientifiques et techniques soulevees par les differentes etapes de la chaine numerique couleur. Cet ouvrage aborde les aspects fondamentaux lies a la colorimetrie et a la physiologie, a la constance et a l'apparence des couleurs. Il traite aussi des aspects plus techniques lies aux capteurs et a la gestion des couleurs sur ecran. Une attention particuliere a ete egalement apportee a la notion de rendu des couleurs en synthese d'images. Au dela de la couleur, un etat de l'art approfondit est aussi mene sur le codage, la compression, la protection et la qualite d'images et de videos couleur.