Today's packet-switched networks are subject to bandwidth fluctuations that cause degradation of the user experience of multimedia services. In order to cope with this problem, HTTP adaptive streaming (HAS) has been proposed in recent years as a video delivery solution for the future Internet and being adopted by an increasing number of streaming services, such as Netflix and Youtube. HAS enables service providers to improve users' quality of experience (QoE) and network resource utilization by adapting the quality of the video stream to the current network conditions. However, the resulting time-varying video quality caused by adaptation introduces a new type of impairment and thus novel QoE research challenges. Despite various recent attempts to investigate these challenges, many fundamental questions regarding HAS perceptual performance are still open. In this paper, the QoE impact of different technical adaptation parameters, including chunk length, switching amplitude, switching frequency, and temporal recency, are investigated. In addition, the influence of content on perceptual quality of these parameters is analyzed. To this end, a large number of adaptation scenarios have been subjectively evaluated in four laboratory experiments and one crowdsourcing study. A statistical analysis of the combined data set reveals results that partly contradict widely held assumptions and provide novel insights in perceptual quality of adapted video sequences, e.g., interaction effects between quality switching direction (up/down) and switching strategy (smooth/abrupt). The large variety of experimental configurations across different studies ensures the consistency and external validity of the presented results that can be utilized for enhancing the perceptual performance of adaptive streaming services.
Given recent demographic changes, adapting the office environments of older knowledge workers to their needs has become increasingly important in supporting an extension of working life. In this paper, we present a case study research of older knowledge workers in Romania, with the goal of gaining bottom-up insights that support the ideation, design, and development of features for a smart work environment. Utilizing a multi-method approach, we combine (1) contextual interviews and observations, (2) an analysis of needs and frictions for deriving insights, (3) an ideation workshop for eliciting potential features, (4) an online survey among experts for evaluating the final feature ideas, and (5) early stage prototyping of selected feature ideas. Following this comprehensive yet efficient approach, we were able to gain a rich understanding of the work realities and contexts of older knowledge workers and to transform that understanding into a concrete set of prioritized feature ideas.
In this work, we identify influencing factors on modality choices of older adults. In detail, we investigated when and why older adults prefer speech over touch interaction and vice versa when interacting with a mobile multimodal health and wellbeing service. We conducted a study with 19 older adults using a mobile application with a duration of three to six weeks. Due to this long duration of the study we were able to gain highly external valid insights as our results are based on real world experiences. We identify additional influencing factors within the areas of user characteristics, contextual factors and perceived system characteristics. We outline the impact of the factors and highlight the importance of several of these factors to enable accessible user interfaces. Our results provide first steps towards a more holistic model of modality choices taking into account interdependencies of different factors.
While the modeling of QoE has made significant advances over the last couple of years, currently existing models still lack an integration of user behavior aspects and user context factors along with the consideration of appropriate temporal scales. Therefore, the goal of this paper is to present a comprehensive QoE and user behavior model providing a framework which allows joining a multitude of existing modeling approaches under the perspectives of service provider benefit, user well-being and technical system performance. In addition, we discuss the role of a broad range of corresponding influence factors, with a specific emphasis on user and context issues, and illustrate our proposal through a series of related use cases.
The optimization as well as exploitation of various aspects of user experience is crucial for future technological innovation and adoption. As a consequence of individualization, industrialization and lifestyle orientation, user experience is becoming more and more a major paradigm in the industry as well as in research & technology organizations. This applies at the level of products (goods, services), at the level of (public) technical infrastructures as well as on the level of human oriented innovation cultures and approaches. Based on several years of experience in applied HCI research the Business Unit Technology Experience within the Innovation Systems Department at the Austrian Institute of Technology (AIT) has been established as a horizontal unit to bridge between innovation in technological infrastructures and the diverse needs of users, costumers or diverse infrastructure contexts. Providing different viewpoints of technology experience and applied HCI thinking is a vehicle to facilitate improved levels of experiential quality.
Crowdsourcing (CS) has evolved into a mature assessment methodology for subjective experiments in diverse scientific fields and in particular for QoE assessment. However, the results acquired for absolute category rating (ACR) scales through CS are often not fully comparable to QoE assessments done in laboratory environments. A possible reason for such differences may be the scale usage heterogeneity problem caused by deviant scale usage of the crowd workers. In this paper, we study different implementations of (quality) rating scales (in terms of design and number of answer categories) in order to identify if certain scales can help to overcome scale usage problems in crowdsourcing. Additionally, training of subjects is well known to enhance result quality for laboratory ACR evaluations. Hence, we analyzed the appropriateness of training conditions to overcome scale usage problems across different samples in crowdsourcing. As major results, we found that filtering of user ratings and different scale designs are not sufficient to overcome scale usage heterogeneity, but training sessions despite their additional costs, enhance result quality in CS and properly counterfeit the identified scale usage heterogeneity problems.
For all Confidential documents (CN, CL, CR): This document contains information that is confidential and proprietary to NGMN Ltd. The information may not be used, disclosed or reproduced without the prior written authorisation of NGMN Ltd., and those so authorised may only use this information for the purpose consistent with the authorisation. For Public documents (P): © 2013 Next Generation Mobile Networks Ltd. All rights reserved. No part of this document may be reproduced or transmitted in any form or by any means without prior written permission from NGMN Ltd.
This paper reports on activities in Study Group 12 of the International Telecommunication Union (ITU-T SG12) to define a new Recommendation on subjective evaluation methods for gaming Quality of Experience (QoE). It first resumes the structure and content of the current draft which has been proposed to ITU-T SG12 in September 2014 and then critically discusses potential gaming content and evaluation methods for inclusion into the upcoming Recommendation. The aim is to start a discussion amongst experts on potential evaluation methods and their limitations, before finalizing a Recommendation. Such a recommendation might in the end be applied by non-expert users, hence wrong decisions in the evaluation design could negatively affect gaming QoE throughout the evaluation.
Cloud-based systems are gaining enormous popularity due to a number of promised benefits, including ease of deployment and administration, scalability and flexibility, and costs savings. However, as more personal and business applications migrate to the Cloud, the service quality becomes an important differentiator between providers. ISPs, Cloud providers and enterprises migrating their services to the Cloud must therefore understand the network requirements to ensure proper end-user Quality of Experience (QoE) in these services. This paper addresses the problem of QoE in Telepresence and Remote Collaboration (TRC) services provided by Microsoft Lync Online (MLO). MLO is a Cloud-based service providing online meeting capabilities including videoconferencing, audio calls, and desktop sharing, and has become the default system for TRC in enterprise scenarios. We present a complete study of the QoE undergone by 44 MLO users in controlled subjective lab tests. The study is performed on three different interactive scenarios running on top of the real MLO Cloud service, additionally shaping the Lync flows at the access network to influence the participants experience. The scenarios include audioconferencing, video-conferencing, and remote collaboration though desktop sharing. By passively monitoring the end-to-end QoS achieved by the Lync flows, and correlating it with the QoE feedbacks provided by the participants, this study permits to better understand the interplays between network performance and QoE in TRC Cloud services. In addition, we provide a network-level characterization of the traffic generated by MLO, as well as an overview on the infrastructure hosting MLO servers.
The chapter discusses the processes of human perception and experiencing, and of quality formation. In this context, definitions of relevant terms are re-visited and adapted to the presented, updated view, and different aspects of research into quality at large and into Quality of Experience are summarized. Using a conceptual model, the quality formation process is analyzed in view of different contexts and tasks, such as taking part in a quality test under controlled conditions, experiencing a video presentation or concert, or exploring a system or device when considering a purchase in a shop. We provide a short overview of different quality assessment methods, and outline related trends in QoE research.
Changing network conditions like bandwidth fluctuations and resulting bad user experience issues (e.g. video freezes) pose severe challenges to Internet video streaming. To address this problem, an increasing number of video services utilizes HTTP adaptive streaming (HAS). HAS enables service providers to improve Quality of Experience (QoE) and resource utilization by incorporating information from different layers. However, these adaptation possibilities of HAS also introduce new perceivable impairments such as the fluctuation of audiovisual quality levels over time, which in turn lead to novel QoE-related research questions. The main contribution of this paper is the formulation of open research questions as well as a thorough systematic user-centric analysis of different quality adaptation dimensions and strategies. The underlying data has been acquired through two crowdsourcing and one lab study. The results provide guidance w.r.t. which encoding dimensions are combined best for the creation of the adaptation set and what type of adaptation strategy should be used. Furthermore it provides insights on the impact of adaptation frequency and the true QoE gain of adaptation over stallings.
HTTP adaptive streaming technology has become widely spread in multimedia services because of its ability to provide adaptation to characteristics of various viewing devices and dynamic network conditions. There are various studies targeting the optimization of adaptation strategy. However, in order to provide an optimal viewing experience to the end-user, it is crucial to get knowledge about the Quality of Experience (QoE) of different adaptation schemes. This paper overviews the state of the art concerning subjective evaluation of adaptive streaming QoE and highlights the challenges and open research questions related to QoE assessment.
Although web applications are well known for performance variations during a session, the impact of network quality fluctuations on user-perceived quality has been neglected so far in QoE research. In this paper, we present the results of two subjective lab experiments which investigated the influence of outages, throughput alternations and increasing/decreasing throughput on Web QoE. We found that interactive applications like Google Maps are heavily impaired by outages, whereas less interactive scenarios like browsing a News Site are more tolerant. Surprisingly, the alternation frequency had no significant effect on subjective quality evaluation. Our results represent a first step towards reliable assessment and modeling of the impact of quality fluctuations on Web QoE.
This work analyses the interaction behaviour of two interlocutors communicating over telephone connections affected by echo-free delay, for conversation tasks yielding different speed and structure. Based on a series of conversation tests, it is shown that transmission delay in a telephone circuit does not only result in a longer time until information is exchanged between the interlocutors, but also alters various characteristics of the conversational course. It was observed that with increasing transmission delay, the realities perceived by the interlocutors increasingly diverge. As a measure of utterance pace, a new conversation surface structure metric, the so-called utterance rhythm (URY), is introduced. Using surface-structure analysis of conversations from different conversation tests, it is shown that peoples' utterance rhythm stays rather constant in close-to-natural conversations, but is considerably affected for scenarios requiring fast interaction and a clear answering structure. At the same time, the quality of the connection is perceived less critically in close-to-natural than in tasks requiring fast interaction, that is, interactive tasks leading to a delay-dependant utterance rhythm. Hence, the conclusion can be drawn that the degree of necessary adaption of the utterance rhythm to a certain delay condition co-determines the extent to which transmission delay impacts the perceived integral quality of a call. (C) 2014 Elsevier B.V. All rights reserved.
This chapter discusses the relation between interactivity and QoE. In this context, a definition of interactivity comprising human-to-human interaction as well as human-to-machine interaction is presented, and a description of a possible instrumentation is given. In terms of quality formation, a mediation layer between quality influence factors and perceived quality features is introduced that allows the inclusion of interactivity-related perception in the quality formation process. A discussion of commonalities and differences between interaction with a system and interaction with one or several other persons via a system identifies the open challenges for reliable and successful measurement of interactivity related aspects and the identification of relationships between these interaction measures and QoE.
Due to the rapidly increasing adoption of broadband Internet services, the user-perceived quality of interactiveWeb and Cloud applications has become a hot topic in QoE research. However, the scientific community still suffers a lack of comparability and reproducibility of results in this field. To address this problem we present the first freely available MOS-labeled dataset forWeb Browsing QoE. The data is based on a subjective lab experiment in which participants had to browse four different websites at different network speeds resulting in different levels of experienced responsiveness. In addition to the dataset, test setup and methodology are briefly discussed.
Crowdsourcing is a popular approach that outsources tasks via the Internet to a large number of users. Commercial crowdsourcing platforms provide a global pool of users employed for performing short and simple online tasks. For quality assessment of multimedia services and applications, crowdsourcing enables new possibilities by moving the subjective test into the crowd resulting in larger diversity of the test subjects, faster turnover of test campaigns, and reduced costs due to low reimbursement costs of the participants. Further, crowdsourcing allows easily addressing additional features like real-life environments. This white paper summarizes the recommendations and best practices for crowdsourced quality assessment of multimedia applications from the Qualinet Task Force on “Crowdsourcing”. The European Network on Quality of Experience in Multimedia Systems and Services Qualinet (COST Action IC 1003, see www.qualinet.eu) established this task force in 2012 which has more than 30 members. The recommendation paper resulted from the experience in designing, implementing, and conducting crowdsourcing experiments as well as the analysis of the crowdsourced user ratings and context data.
Real-world usage of web browsing differs considerably from typically employed laboratory tests on single or multiple page views. Especially when using mobile devices such as smartphones or tablets, several sources of distraction are prevalent while browsing. In the mobile scenario, users are exposed to distracting factors like other people, traffic, announcements etc., whereas at home parallel use of TV or radio services might distract the user. To assess the impact due to context and distractions, this paper presents two studies of web browsing Quality of Experience. The studied factors of possible influence on QoE include one specific browsing task, the environment and a dedicated distraction task. The results show that, in contrast to ratings for multimedia sessions containing audio, video or speech, QoE ratings for web browsing are not affected by the considered contexts or distractions. However, it was found that the primary task for the browsing session had a significant influence on the QoE ratings regardless of the environment or distracting task.
Since its introduction a few years ago, the concept of `Crowdsourcing' has been heralded as highly attractive alternative approach towards evaluating the Quality of Experience (QoE) of networked multimedia services. The main reason is that, in comparison to traditional laboratory-based subjective quality testing, crowd-based QoE assessment over the Internet promises to be not only much more cost-effective (no lab facilities required, less cost per subject) but also much faster in terms of shorter campaign setup and turnaround times. However, the reliability of remote test subjects and consequently, the trustworthiness of study results is still an issue that prevents the widespread adoption of crowd-based QoE testing. Various ideas for improving user rating reliability and test efficiency have been proposed, with the majority of them relying on a posteriori analysis of results. However, such methods introduce a major lag that significantly affects efficiency of campaign execution. In this paper we address these shortcomings by introducing in momento methods for crowdsourced video QoE assessment which yield improvements of results reliability by factor two and campaign execution efficiency by factor ten. The proposed in momento methods are applicable to existing crowd-based QoE testing approaches and suitable for a variety of service scenarios.
Article QoE Analysis for Interactive Internet Applications in the Presence of Delay was published on December 1, 2014 in the journal PIK - Praxis der Informationsverarbeitung und Kommunikation (volume 37, issue 4).