The paper presents a series of three new video quality model standards for the assessment of sequences of up to UHD/4K resolution. They were developed in a competition within the International Telecommunication Union (ITU-T), Study Group 12, in collaboration with the Video Quality Experts Group (VQEG), over a period of more than two years. A large video quality test set with a total of 26 individual databases was created, with 13 used for training and 13 for validation and selection of the winning models. For each database, video quality laboratory tests were run with at least 24 subjects each. The 5-point Absolute Category Rating (ACR) scale was used for rating, calculating Mean Opinion Scores (MOS) as ground-truth. To represent today's commonly applied HTTP-based adaptive streaming context, the test sequences comprise a variety of encoding settings, bitrates, resolutions and framerates for the three codecs H.264/AVC, H.265/HEVC and VP9, applied to a wide range of source sequences of around 8 s duration. Processing was carried out with an FFmpeg-based processing chain developed specifically for the competition, and via upload and encoding through exemplary online streaming services. The resulting data represents the largest, lab-test-based dataset used for video quality model development to date, with a total of around 5,000 test sequences. The paper addresses the three models ultimately standardized in the P.1204 Recommendation series, resulting in different model types and for different applications: (i) Rec. P.1204.3, no-reference bitstream-based, with access to encoded bitstream information; (ii) P.1204.4, pixel-based, using information from the reference and the processed video, and (iii) P.1204.5, no-reference hybrid, using both bitstream-and pixel-information without knowledge of the reference. The paper outlines the development process and provides holistic details about the statistical evaluation, test databases, model algorithms and validation results, as well as a performance comparison with state-of-the-art models.
With the increasing requirement of users to view high-quality videos with a constrained bandwidth, typically realized using HTTP-based adaptive streaming, it becomes more and more important to determine the quality of the encoded videos accurately, to assess and possibly optimize the overall streaming quality. In this paper, we describe a bitstream-based no-reference video quality model developed as part of the latest model-development competition conducted by ITU-T Study Group 12 and the Video Quality Experts Group (VQEG), “P.NATS Phase 2”. It is now part of the new P.1204 series of Recommendations as P.1204.3. It can be applied to bitstreams encoded with H.264/AVC, HEVC and VP9, using various encoding options, including resolution, bitrate, framerate and typical encoder settings such as number of passes, rate control variants and speeds. The proposed model follows an ensemble-modelling-inspired approach with weighted parametric and machine-learning parts to efficiently leverage the performance of both approaches. The paper provides details about the general approach to modelling, the features used and the final feature aggregation. The model creates per-segment and per-second video quality scores on the 5-point Absolute Category Rating scale, and is applicable to segments of 5–10 seconds duration. It covers both PC/TV and mobile/tablet viewing scenarios. We outline the databases on which the model was trained and validated as part of the competition, and perform an additional evaluation using a total of four independently created databases, where resolutions varied from 360p to 2160p, and frame rates from 15–60fps, using realistic coding and bitrate settings. We found that the model performs well on the independent dataset, with a Pearson correlation of 0.942 and an RMSE of 0.42. We also provide an open-source reference implementation of the described P.1204.3 model, as well as the multi-codec bitstream parser required to extract the input data, which is not part of the standard.
Adaptive streaming is fast becoming the most widely used method for video delivery to the end users over the internet. The ITU-T P.1203 standard is the first standardized quality of experience model for audiovisual HTTP-based adaptive streaming. This recommendation has been trained and validated for H.264 and resolutions up to and including full-HD. The paper provides an extension for the existing standardized short-term video quality model mode 0 for new codecs i.e., H.265, VP9 and AV1 and resolutions larger than full-HD (e.g. UHD-1). The extension is based on two subjective video quality tests. In the tests, in total 13 different source contents of 10 seconds each were used. These sources were encoded with resolutions ranging from 360p to 2160p and various quality levels using the H.265, VP9 and AV1 codecs. The subjective results from the two tests were then used to derive a mapping/correction function for P.1203.1 to handle new codecs and resolutions. It should be noted that the standardized model was not re-trained with the new subjective data, instead only a mapping/correction function was derived from the two subjective test results so as to extend the existing standard to the new codecs and resolutions.
This paper describes an open dataset and software for ITU-T Ree. P.1203. As the first standardized Quality of Experience model for audiovisual HTTP Adaptive Streaming (HAS), it has been extensively trained and validated on over a thousand audiovisual sequences containing HAS-typical effects (such as stalling, coding artifacts, quality switches). Our dataset comprises four of the 30 official subjective databases at a bitstream feature level. The paper also includes subjective results and the model performance. Our software for the standard was made available to the public, too, and it is used for all the analyses presented. Among other previously unpublished details, we show the significant performance improvements of using bitstream-based models over metadata-based ones for video quality analysis, and the robustness of combining classical models with machine-learning-based approaches for estimating user QoE.
The paper presents the scalable video quality model part of the P.1203 Recommendation series, developed in a competition within ITU-T Study Group 12 previously referred to as P.NATS. It provides integral quality predictions for 1 up to 5 min long media sessions for HTTP Adaptive Streaming (HAS) with up to HD video resolution. The model is available in four modes of operation for different levels of media-related bitstream information, reflecting different types of encryption of the media stream. The video quality model presented in this paper delivers short-term video quality estimates that serve as input to the integration component of the P.1203 model. The scalable approach consists in the usage of the same components for spatial and temporal scaling degradations across all modes. The third component of the model addresses video coding artifacts. To this aim, a single model parameter is introduced that can be derived from different types of bitstream input information. Depending on the complexity of the available input, one of four scaling-levels of the model is applied. The paper presents the different novelties of the model and scientific choices made during its development, the test design, and an analysis of the model performance across the different modes.
In this paper, a novel method to detect scene changes in encrypted video streams is presented. Typically, in IPTV systems, the media stream is transmitted in encrypted form, and therefore the only available information to determine the scene changes are the packet headers which transport the video signal. Thus, the proposed method estimates the size and the type of each picture of the video sequence by extracting information from the packet headers. Then, based on the GOP structure, a set of rules are determined to predict changes of frame sizes which are indicative of scene changes in the video sequences. Furthermore, the application of the proposed method in the recently standardized ITU-T Recommendation P.1201.2 for no-reference audio-visual quality assessment for IPTV-grade services is presented to highlight how such method could be deployed. Finally, the proposed method is evaluated on a large set of video databases to demonstrate the validity of the proposed method.
A parametric packet-based model has been created to estimate user perceived audiovisual quality of Internet Protocol Television (IPTV) services. It is divided into three modules, for audio, video and audiovisual quality. The model is applicable to the quality monitoring of encrypted and non-encrypted audiovisual streams. Typical audio and video degradations for IPTV are covered for Standard Definition (SD) and High Definition (HD) video formats. The model supports the H.264 video codec and the audio codecs MPEG-I Layer II, MPEG-2 AAC-LC, MPEG-4 HE-AACv2 and AC3. It handles various types of IP-network layer transmission errors. The model was developed and validated using a large database of subjective tests. The underlying concept is based on an impairment factor approach, which enables detection of how users build their individual judgment of quality of a given audiovisual signal. Each impairment factor captures the perceived quality impact of a possible degradation and therefore enables diagnostic analysis of quality problems. The model shows high performance results, both in terms of Pearson's Correlation coefficient (r) and Root-Mean-Square-Error (RMSE). The model is standardized as ITU-T Recommendation P.1201.2, the higher resolution (IPTV and Video on Demand (VoD)) algorithm of Recommendation P.1201.
In this paper, a novel method for predicting the visibility of packet losses in SD and HD H.264/AVC video sequences and modeling their impact on perceived quality is proposed. Based on the findings of a new subjective experiment it is initially shown that the classification of packet loss visibility in a binary fashion is not sufficient to model the perceptual degradations caused by the transmission errors. The proposed no-reference algorithm extracts a set of features from the video bit-stream to account for the spatial and temporal characteristics of the video content and the induced distortion due to the network impairments. Subsequently, the visibility of packet losses is predicted in a continuous fashion using Support Vector Regression. Finally, a no-reference bit-stream based video quality assessment model that explicitly employs the predicted packet loss visibility estimates is presented. The evaluation of the proposed model demonstrates that the use of continuous estimates for the visibility of packet losses improves the performance of the video quality assessment model.
This article provides a tutorial overview of current approaches for monitoring the quality perceived by users of IP-based audiovisual media services. The article addresses both mobile and fixed network services such as mobile TV or Internet Protocol TV (IPTV). It reviews the different quality models that exploit packet- header-, bit stream-, or signal-information for providing audio, video, and audiovisual quality estimates, respectively. It describes how these models can be applied for real-life monitoring, and how they can be adapted to reflect the information available at the given measurement point. An outlook gives insight into emerging trends for near- and mid-term future requirements and solutions.
Method for evaluating the quality of a video signal transmitted in the receiver side, the method comprising the steps of: a) capturing the input bit stream and deliver it to an analyzer of video bitstream; b) extracting at least one feature or set of features from the bit stream input video captured by the bitstream analyzer; c) delivering the extracted feature or set of features to an estimation module visibility packet loss; d) determining, by the estimation module visibility packet loss, using the extracted features supplied bit stream of video, the continuous probability of visibility for each event packet loss occurs within a range of specific time; e) use the continuous probability of visibility of packet loss, determined by the estimation module visibility packet loss, as a weighting factor of at least one feature or set of features extracted from the flow of video bits to calculate an estimate of the overall quality, Q, of the transmitted video sequence; wherein step (d) uses at least one characteristic of the bitstream from the group comprising: frame type, average magnitude of vector motion (AvgMv), average difference vector motion (AvgMvDiff), residual energy (ResEnergy), maximum number of partitions (MaxPartNr), number of non-decodable macroblock (LostMbs), motion vector information (mv), macroblock type (mb type); and wherein step (e) combines the estimation of visibility (V) of packet loss to the determined magnitude of distortion (EstErr) and the calculated total number of bad pixels due to packet loss (ErrProp ).
In this paper, a no reference bit stream model for quality assessment of SD and HD H.264/AVC video sequences based on packet loss visibility is proposed. The method considers the impact of network impairments on human perception and uses the visibility of packet losses to predict objective scores. Also, a new subjective experiment has been designed to provide insight into the perceptual effect of degradations caused by transmission errors. The proposed algorithm extracts a set of features from the received bit stream. Then, the visibility of each packet loss event is determined by classifying the extracted features using a Support Vector Machines classifier. Finally, analytical expressions are developed to account for visual degradation due to compression and channel induced distortion based on the outcome of the visibility classifier. The evaluation demonstrates the validity of the proposed method.
The paper presents a parameter-based model for predicting the perceived quality of transmitted video for IPTV applications. The core model we derived can be applied both to service monitoring and network or service planning. In its current form, the model covers H.264 and MPEG-2 coded video (standard and high definition) transmitted over IP-links. The model includes factors like the coding bit-rate, the packet loss percentage and the type of packet loss handling used by the codec. The paper provides an overview of the model, of its integration into a multimedia model predicting audiovisual quality, and of its application to service monitoring. A performance analysis is presented showing a high correlation with the results of different subjective video quality perception tests. An outlook highlights future model extensions.