Cross-laboratory subjective quality studies rely on Mean Opinion Scores (MOS) as ground truth, yet systematic differences between laboratories persist even after alignment. The VQEG Multimedia Phase 2 study showed that MOS is relative rather than absolute, with individual subject differences being the second-most important source of variability - but these were never quantified. We use SUREAL to recover subject-level bias parameters and decompose variance through pooled vs. perlaboratory analyses. Results show that 85% of bias variance reflects genuine subject differences, while 15% is attributable to lab-level systematic effects. We also find that content-dependent laboratory effects are strongest for animated and music content. The substantial within-country variation suggests that laboratory-specific factors may mask cultural influences rather than negate them. These findings argue for reporting subject-level parameter distributions alongside MOS.
Demand for streaming services, including satellite, continues to exhibit unprecedented growth. As subscribers increasingly demand seamless streaming, satellite Internet Service Providers (ISPs) find themselves at the crossroads of technological advancements and rising customer expectations. To stay relevant and competitive, these ISPs must ensure they deliver optimal streaming video quality, a key determinant of user satisfaction. Towards this end, it is important to have accurate Quality of Experience (QoE) prediction models in place. These models can be used to optimize streaming protocols and ensure uninterrupted viewing experiences. However, achieving robust performance by these models requires extensive data sets labeled by subjective opinion scores on videos impaired by diverse playback disruptions. To bridge this data gap, we introduced the LIVE-Viasat Real-World Satellite QoE Database. This database consists of 179 videos recorded from real-world streaming services affected by various authentic distortion patterns. We also conducted a comprehensive subjective study involving 54 participants, who contributed both continuous-time opinion scores and endpoint (retrospective) QoE scores. Our analysis sheds light on various determinants influencing subjective QoE, such as stall events, spatial resolutions, bitrate, and certain transmission parameters. We demonstrate the usefulness of this unique new resource by evaluating the efficacy of prevalent QoE-prediction models on it. We also developed and introduced a new model, called SatQA, that maps network parameters to human judgments of perceptual QoE. This tool can be used to help Internet Service Providers (ISPs) optimize video quality streamed from satellites to millions of customers. Unlike prior methods, SatQA achieves accurate predictions of QoE using only transmission parameters, without requiring access to pixel data or video-specific metadata. Moreover, it is trained on human MOS in the new LIVE-Viasat dataset. Experiments show that SatQA achieves high QoE prediction accuracy and reliability as compared to prior leading algorithms. The publicly available database can be found at https://live.ece.utexas.edu/research/LIVE_Viasat_Real_World_Satellite_QoE/index.html.
Reaching close-to-optimal bandwidth utilization in Dynamic Adaptive Streaming over HTTP (DASH) systems can, in theory, be achieved with a small discrete set of bit rate representations. This includes typical bit rate ladders used in state-of-the-art DASH systems. In practice, however, we demonstrate that bandwidth utilization, and consequently the Quality of Experience (QoE), can be improved by offering a continuous set of bit rate representations, i.e., a continuous bit rate slide (COBIRAS). Moreover, we find that the buffer fill behavior of different standard adaptive bit rate (ABR) algorithms is sub-optimal in terms of bandwidth utilization. To overcome this issue, we leverage COBIRAS’ flexibility to request segments with any arbitrary bit rate and propose a novel ABR algorithm MinOff , which helps maximizing bandwidth utilization by minimizing download off-phases during streaming. To avoid extensive storage requirements with COBIRAS and to demonstrate the feasibility of our approach, we design and implement a proof-of-concept DASH system for video streaming that relies on just-in-time encoding ( JITE ), which reduces storage consumption on the DASH server. Finally, we conduct a performance evaluation on our testbed and compare a state-of-the-art DASH system with few bit rate representations and our JITE DASH system, which can offer a continuous bit rate slide, in terms of bandwidth utilization and video QoE for different ABR algorithms.
This paper uses a crowdsourced dataset of online video streaming sessions to investigate opportunities to reduce the power consumption while considering QoE. For this, we base our work on prior studies which model both the end-user's QoE and the end-user device's power consumption with the help of high-level video features such as the bitrate, the frame rate, and the resolution. On top of existing research, which focused on reducing the power consumption at the same QoE optimizing video parameters, we investigate potential power savings by other means such as using a different playback device, a different codec, or a predefined maximum quality level. We find that based on the power consumption of the streaming sessions from the crowdsourcing dataset, devices could save more than 55% of power if all participants adhere to low-power settings.
Audiovisual asynchrony is a significant factor im-pacting the Quality of Experience (QoE), especially for interactive communication like video conferencing. In this paper, we propose a client-side approach to predict the delay between an audio and a video signal, using only the media signals from both streams. Features are extracted from the video and audio stream, respectively, and analyzed using a cross-correlation approach to determine the actual delay. Our approach predicts the delay with an accuracy of over 80% in a time frame of ±1s. We further highlight the potential drawbacks of using a cross-correlation-based analysis and propose different solutions for practical implementations of a delay-based QoE metric.
This paper presents two challenges associated with using the ITU-T P.1203 standard for video quality monitoring in practice. We discuss the issue of unavailable data on certain browsers/platforms and the lack of information within newly developed data formats like Common Media Client Data. We also re-trained the coefficients of the P.1203.1 video model for newer codecs, and published a completely new model derived from the P.1204.3 bitstream model.
In the past decade, we have witnessed an enormous growth in the demand for online video services. Recent studies estimate that nowadays, more than 1% of the global greenhouse gas emissions can be attributed to the production and use of devices performing online video tasks. As such, research on the true power consumption of devices and their energy efficiency during video streaming is highly important for a sustainable use of this technology. At the same time, over-the-top providers strive to offer high-quality streaming experiences to satisfy user expectations. Here, energy consumption and QoE partly depend on the same system parameters. Hence, a joint view is needed for their evaluation. In this paper, we perform a first analysis of both end-user power efficiency and Quality of Experience of a video streaming service. We take a crowdsourced dataset comprising 447,000 streaming events from YouTube and estimate both the power consumption and perceived quality. The power consumption is modeled based on previous work which we extended towards predicting the power usage of different devices and codecs. The user-perceived QoE is estimated using a standardized model. Our results indicate that an intelligent choice of streaming parameters can optimize both the QoE and the power efficiency of the end user device. Further, the paper discusses limitations of the approach and identifies directions for future research.
Welcome to the fourth column on the ACM SIGMM Records from the Video Quality Experts Group (VQEG). During the last VQEG plenary meeting (14-18 Dec. 2020) various interesting discussions arose regarding new topics not addressed up to then by VQEG groups, which led to launching three new sub-projects and a new project related to: 1) clarifying the computation of spatial and temporal information (SI and TI), 2) including video quality metrics as metadata in compressed bitstreams, 3) Quality of Experience (QoE) metrics for live video streaming applications, and 4) providing guidelines on implementing objective video quality metrics to the video compression community. The following sections provide more details about these new activities and try to encourage interested readers to follow and get involved in any of them by subscribing to the corresponding reflectors.
This article summarizes definitions of crowdsourcing in the context of network and QoE measurements as provided in White Paper on Crowdsourced Network and QoE Measurements - Definitions, Use Cases and Challenges (2020). Tobias Hoßfeld and Stefan Wunderer, eds., Würzburg, Germany, March 2020. doi: 10.25972/OPUS-20232.
Spatial Information (SI) and Temporal Information (TI) are frequently-used metrics to classify the spatiotemporal complexity of video content. However, they are mostly used on original video sources, and their impact on actual encoding efficiency is not known. In this paper, we propose a method to determine the compressibility of video sources, that is, how good video quality can be under a given bitrate constraint. We show how various aggregations of SI and TI correlate with compressibility scores obtained from a public dataset of H.264/HEVCN P9 content. We observe that the minimum TI value as well as an existing criticality metric from the literature are good indicators for compressibility, as judged by subjective ratings as well as VMAF and P.1204.3 objective scores.
HTTP Adaptive Streaming (HAS) is the de-facto standard for video delivery over the Internet. It enables dynamic adaptation of video quality by splitting a video into small segments and providing multiple quality levels per segment. So far, HAS services typically utilize a fixed segment duration. This reduces the encoding and streaming variability and thus allows a faster encoding of the video content and a reduced prediction complexity for adaptive bit rate algorithms. Due to the content-agnostic placement of I-frames at the beginning of each segment, additional encoding overhead is introduced. In order to mitigate this overhead, variable segment durations, which take encoder placed I-frames into account, have been proposed recently. Hence, a lower number of I-frames is needed, thus achieving a lower video bitrate without quality degradation. While several proposals exploiting variable segment durations exist, no comparative study highlighting the impact of this technique on coding efficiency and adaptive streaming performance has been conducted yet. This paper conducts such a holistic comparison within the adaptive video streaming eco-system. Firstly, it provides a broad investigation of video encoding efficiency for variable segment durations. Secondly, a measurement study evaluates the impact of segment duration variability on the performance of HAS using three adaptation heuristics and the dash.js reference implementation. Our results show that variable segment durations increased the Quality of Experience for 54% of the evaluated streaming sessions, while reducing the overall bitrate by 7% on average.
This white paper is the outcome of the Würzburg seminar on "Crowdsourced Network and QoE Measurements" which took place from 25-26 September 2019 in Würzburg, Germany. International experts were invited from industry and academia. They are well known in their communities, having different backgrounds in crowdsourcing, mobile networks, network measurements, network performance, Quality of Service (QoS), and Quality of Experience (QoE). The discussions in the seminar focused on how crowdsourcing will support vendors, operators, and regulators to determine the Quality of Experience in new 5G networks that enable various new applications and network architectures. As a result of the discussions, the need for a white paper manifested, with the goal of providing a scientific discussion of the terms "crowdsourced network measurements" and "crowdsourced QoE measurements", describing relevant use cases for such crowdsourced data, and its underlying challenges. During the seminar, those main topics were identified, intensively discussed in break-out groups, and brought back into the plenum several times. The outcome of the seminar is this white paper at hand which is - to our knowledge - the first one covering the topic of crowdsourced network and QoE measurements.
In recent years, with the introduction of powerful HMDs such as Oculus Rift, HTC Vive Pro, the QoE that can be achieved with VR/360° videos has increased substantially. Unfortunately, no standardized guidelines, methodologies and protocols exist for conducting and evaluating the quality of 360° videos in tests with human test subjects. In this paper, we present a set of test protocols for the evaluation of quality of 360° videos using HMDs. To this aim, we review the state-of-the-art with respect to the assessment of 360° videos summarizes their results. Also, we summarize the methodological approaches and results taken for different subjective experiments at our lab under different contextual conditions. In the first two experiments 1a and 1b, the performance of two different subjective test methods, Double-Stimulus Impairment Scale (DSIS) and Modified Absolute Category Rating (M-ACR) was compared under different contextual conditions. In experiment 2, the performance of three different subjective test methods, DSIS, M-ACR and Absolute Category Rating (ACR) was compared this time without varying the contextual conditions. Building on the reliability and general applicability of the procedure across different tests, a methodological framework for 360° video quality assessment is presented in this paper. Besides video or media quality judgments, the procedure comprises the assessment of presence and simulator sickness, for which different methods were compared. Further, the accompanying head-rotation data can be used to analyze both content- and quality-related behavioural viewing aspects. Based on the results, the implications of different contextual settings are discussed.
The paper presents a series of three new video quality model standards for the assessment of sequences of up to UHD/4K resolution. They were developed in a competition within the International Telecommunication Union (ITU-T), Study Group 12, in collaboration with the Video Quality Experts Group (VQEG), over a period of more than two years. A large video quality test set with a total of 26 individual databases was created, with 13 used for training and 13 for validation and selection of the winning models. For each database, video quality laboratory tests were run with at least 24 subjects each. The 5-point Absolute Category Rating (ACR) scale was used for rating, calculating Mean Opinion Scores (MOS) as ground-truth. To represent today's commonly applied HTTP-based adaptive streaming context, the test sequences comprise a variety of encoding settings, bitrates, resolutions and framerates for the three codecs H.264/AVC, H.265/HEVC and VP9, applied to a wide range of source sequences of around 8 s duration. Processing was carried out with an FFmpeg-based processing chain developed specifically for the competition, and via upload and encoding through exemplary online streaming services. The resulting data represents the largest, lab-test-based dataset used for video quality model development to date, with a total of around 5,000 test sequences. The paper addresses the three models ultimately standardized in the P.1204 Recommendation series, resulting in different model types and for different applications: (i) Rec. P.1204.3, no-reference bitstream-based, with access to encoded bitstream information; (ii) P.1204.4, pixel-based, using information from the reference and the processed video, and (iii) P.1204.5, no-reference hybrid, using both bitstream-and pixel-information without knowledge of the reference. The paper outlines the development process and provides holistic details about the statistical evaluation, test databases, model algorithms and validation results, as well as a performance comparison with state-of-the-art models.
As video streaming accounts for the majority of Internet traffic, monitoring its quality is of importance to both Over the Top (OTT) providers as well as Internet Service Providers (ISPs). While OTTs have access to their own analytics data with detailed information, ISPs often have to rely on automated network probes for estimating streaming quality, and likewise, academic researchers have no information on actual customer behavior. In this paper, we present first results from a large-scale crowdsourcing study in which three major video streaming OTTs were compared across five major national ISPs in Germany. We not only look at streaming performance in terms of loading times and stalling, but also customer behavior (e.g., user engagement) and Quality of Experience based on the ITUT P.1203 QoE model. We used a browser extension to evaluate the streaming quality and to passively collect anonymous OTT usage information based on explicit user consent. Our data comprises over 400,000 video playbacks from more than 2,000 users, collected throughout the entire year of 2019. The results show differences in how customers use the video services, how the content is watched, how the network influences video streaming QoE, and how user engagement varies by service. Hence, the crowdsourcing paradigm is a viable approach for third parties to obtain streaming QoE insights from OTTs.
Users may experience symptoms of simulator sickness while watching 360 ° /VR videos with Head-Mounted Displays (HMDs). At present, practically no solution exists that can efficiently eradicate the symptoms of simulator sickness from virtual environments. Therefore, in the absence of a solution, it is required to at least quantify the amount of sickness. In this paper, we present initial work on our Simulator Sickness Model SiSiMo including a first component to predict simulator sickness scores over time. Using linear regression of short term scores already shows promising performance for predicting the scores collected from a number of user tests.
With the increasing requirement of users to view high-quality videos with a constrained bandwidth, typically realized using HTTP-based adaptive streaming, it becomes more and more important to determine the quality of the encoded videos accurately, to assess and possibly optimize the overall streaming quality. In this paper, we describe a bitstream-based no-reference video quality model developed as part of the latest model-development competition conducted by ITU-T Study Group 12 and the Video Quality Experts Group (VQEG), “P.NATS Phase 2”. It is now part of the new P.1204 series of Recommendations as P.1204.3. It can be applied to bitstreams encoded with H.264/AVC, HEVC and VP9, using various encoding options, including resolution, bitrate, framerate and typical encoder settings such as number of passes, rate control variants and speeds. The proposed model follows an ensemble-modelling-inspired approach with weighted parametric and machine-learning parts to efficiently leverage the performance of both approaches. The paper provides details about the general approach to modelling, the features used and the final feature aggregation. The model creates per-segment and per-second video quality scores on the 5-point Absolute Category Rating scale, and is applicable to segments of 5–10 seconds duration. It covers both PC/TV and mobile/tablet viewing scenarios. We outline the databases on which the model was trained and validated as part of the competition, and perform an additional evaluation using a total of four independently created databases, where resolutions varied from 360p to 2160p, and frame rates from 15–60fps, using realistic coding and bitrate settings. We found that the model performs well on the independent dataset, with a Pearson correlation of 0.942 and an RMSE of 0.42. We also provide an open-source reference implementation of the described P.1204.3 model, as well as the multi-codec bitstream parser required to extract the input data, which is not part of the standard.
The test methods recommended by the International Telecommunication Union (ITU) for assessing 2D video quality are often used for evaluating omnidirectional / 360° videos. In this paper, we compare the performance of three different test methods, Absolute Category Rating (ACR), a modified version of ACR (M-ACR) with double presentation of the test stimulus, and DSIS (Double Stimulus Impairment Scale), based on the statistical reliability, assessment time and simulator sickness. Different settings were used for HEVC encoding of five 360° source videos of 10 s duration. Results indicate that DSIS is statistically more reliable with higher resolving power, followed by M-ACR and ACR. We found that simulator sickness increases with time, but can be reduced by taking breaks in-between the test sessions. The results for simulator sickness are compared across test methods and with similar tests conducted under different contextual conditions. We also recorded and analyzed the exploration behaviour of the users. Apart from the methodological findings, the test results provide insights into video quality for different resolution and encoding settings (“bitrate ladders”). These may be useful for choosing appropriate representations in the context of HTTP-based adaptive streaming in case of full-frame streaming.