This research investigates the influence of environmental regulation on subjective evaluations of video quality within the Quality of Experience (QoE) paradigm. This work presents a supplementary experiment conducted in a controlled laboratory setting, building on our previous crowdsourcing studies carried out in uncontrolled, web-based conditions using the Prolific platform. Both tests utilized the identical crowdsourcing platform and complied with the International Telecommunication Union Telecommunication (ITU-T) P.910 Recommendations, ensuring external validity and methodological consistency. Participants assessed a collection of processed video sequences (PVS) comprising 46 distinct video clips utilizing the 5-point Absolute Category Rating (ACR) scale, while their response times were documented in milliseconds as measures of cognitive exertion and decision delay. The comparison analysis employs nonparametric tests (Mann-Whitney U and Kolmogorov-Smirnov) and a hierarchical Linear Mixed-Effects Model (LMM) to examine disparities in reaction time distributions, rating consistency, and the incidence of outliers across both environments. The results indicate that controlled settings produce statistically significantly less response variability and enhanced data reliability, whereas uncontrolled settings encompass greater external diversity and real-world unpredictability. These findings offer significant insights into the balance between experimental control and external validity in crowdsourced video quality assessment, advancing the development of scalable approaches for Quality of Experience research.
The video streaming industry is growing. There is demand for high quality videos. These videos are stream to the consumers with a promising quality and low latency. There are various methods to measure the video quality of experience (QoE) in a streaming environment. The main goal of this paper is to provide an overview of methods and techniques to measure the QoE in adaptive streaming domain. This paper provide overview of metrics and QoE models which asses the video quality in streaming. This paper also discusses the dataset exist in video streaming. This paper highlights the challenges and future strategies that should be considered building models for assessing the video quality in adaptive streaming.
Video-centric applications have seen significant growth in recent years with HTTP Adaptive Streaming (HAS) becoming a widely adopted method for video delivery. Recently, low-latency (LL) adaptive bitrate (ABR) algorithms have recently been proposed to reduce the end-to-end delay in HTTP adaptive streaming. This study investigates whether low-latency adaptive bitrate (LL-ABR) algorithms, in their effort to reduce delay, also compromise video quality. To this end, this study presents both objective and subjective evaluation of user experience with traditional DASH and low-latency ABR algorithms. The study employs crowdsourcing to evaluate user-perceived video quality in low-latency MPEG-DASH streaming, with a particular focus on the impact of short segment durations. We also investigate the extent to which quantitative QoE (Quality of Experience) metrics correspond to the subjective evaluation results. Results show that the Dynamic algorithm outperforms the low-latency algorithms, achieving higher stability and perceptual quality. Among low-latency methods, Low-on-Latency (LOL+) demonstrates superior QoE compared to Learn2Adapt-LowLatency (L2A-LL), which tends to sacrifice visual consistency for latency gains. The findings emphasize the importance of integrating subjective evaluation into the design of ABR algorithms and highlight the need for user-centric and perceptually aware optimization strategies in low-latency streaming systems. Our results show that the subjective scores do not always align with objective performance metrics. The viewers are found to be sensitive to complex or high-motion content, where maintaining a consistent user experience becomes challenging despite favorable objective performance metrics.
The use of video streaming is constantly increasing. High-resolution video requires resources on both the sender and the receiver side. Many compression techniques can be utilized to compress the video and simultaneously maintain quality. The main goal of this paper is to provide an overview of video streaming and QoE. This paper describes the basic concepts and discusses existing methodologies to measure QoE. Subjective, objective, and video compression technologies are discussed. This review paper gathers the codec implementation developed by MPEG, Google, and Apple. This paper outlines the challenges and future research directions that should be considered in the measurement and assessment of the quality of experience for video services.
Counting and detecting occluded faces in a crowd is a challenging task in computer vision. In this paper, we propose a new approach to face detection-based crowd estimation under significant occlusion and head posture variations. Most state-of-the-art face detectors cannot detect excessively occluded faces. To address the problem, an improved approach to training various detectors is described. To obtain a reasonable evaluation of our solution, we trained and tested the model on our substantially occluded data set. The dataset contains images with up to 90 degrees out-of-plane rotation and faces with 25%, 50%, and 75% occlusion levels. In this study, we trained the proposed model on 48,000 images obtained from our dataset consisting of 19 crowd scenes. To evaluate the model, we used 109 images with face counts ranging from 21 to 905 and with an average of 145 individuals per image. Detecting faces in crowded scenes with the underlying challenges cannot be addressed using a single face detection method. Therefore, a robust method for counting visible faces in a crowd is proposed by combining different traditional machine learning and convolutional neural network algorithms. Utilizing a network based on the VGGNet architecture, the proposed algorithm outperforms various state-of-the-art algorithms in detecting faces ‘in-the-wild’. In addition, the performance of the proposed approach is evaluated on publicly available datasets containing in-plane/out-of-plane rotation images as well as images with various lighting changes. The proposed approach achieved similar or higher accuracy.
Solutions for emotion recognition are becoming more popular every year, especially with the growth of computer vision. In this paper, classification of emotions is conducted based on images processed with convolutional neural networks (CNNs). Several models are proposed, both custom and transfer learning types. Furthermore, combinations of them as ensembles, alongside various methods of dataset modification, are presented. In the beginning, the models were tested on the original FER2013 dataset. Then, dataset filtering and augmentation were introduced, and the models were retrained accordingly. Two methods of emotion classification were examined: a multi-class classification, and a binary classification. In the former approach, the model returns the probability for each class. In the latter, separate models for each single class are prepared, together with an adequate dataset based on FER2013. Each model recognizes a single emotion from the others. The obtained results and a comparison of the applied methods across different models is presented and discussed.
Video streaming is growing exponentially. High-resolution videos require high bandwidth to transport the videos over the network. There is a great demand for compression technologies to compress video and maintain quality. Video codecs are used to encode and decode video streams. These codecs have been developed by MPEG, Google, Microsoft, and Apple Inc. The goal of this research is to develop a technology that will realize contribution transmission through connecting the latest methods generation of single and multi-way video encoding with the new protocols that will provide transmission reliability and keep low latency. A literature review will be carried out. The literature covers different video codecs and transmission techniques and the methods used to evaluate the quality of those techniques and codecs. Based on the literature review, the theoretical framework will be formed, and a video encoding method prototype will be developed. The developed method will be for one-way and multi-way software that will automatically optimize the settings of the video codec to set its operating conditions at optimal. The new transmission software will use newer codecs, such as H265/HEVC, VP9, and AV1, MPEG5, which will allow additional reduction of the bit stream and deliver secure, reliable, and quality video with low latency.
In the five years between 2017 and 2022, IP video traffic tripled, according to Cisco. User-Generated Content (UGC) is mainly responsible for user-generated IP video traffic. The development of widely accessible knowledge and affordable equipment makes it possible to produce UGCs of quality that is practically indistinguishable from professional content, although at the beginning of UGC creation, this content was frequently characterized by amateur acquisition conditions and unprofessional processing. In this research, we focus only on UGC content, whose quality is obviously different from that of professional content. For the purpose of this paper, we refer to "in the wild" as a closely related idea to the general idea of UGC, which is its particular case. Studies on UGC recognition are scarce. According to research in the literature, there are currently no real operational algorithms that distinguish UGC content from other content. In this study, we demonstrate that the XGBoost machine learning algorithm (Extreme Gradient Boosting) can be used to develop a novel objective "in the wild" video content recognition model. The final model is trained and tested using video sequence databases with professional content and "in the wild" content. We have achieved a 0.916 accuracy value for our model. Due to the comparatively high accuracy of the model operation, a free version of its implementation is made accessible to the research community. It is provided via an easy-to-use Python package installable with Pip Installs Packages (pip).
According to Cisco, we are facing a three-fold increase in IP traffic in five years, ranging from 2017 to 2022. IP video traffic generated by users is largely related to user-generated content (UGC). Although at the beginning of UGC creation, this content was often characterised by amateur acquisition conditions and unprofessional processing, the development of widely available knowledge and affordable equipment allows one to create UGC of a quality practically indistinguishable from professional content. Since some UGC content is indistinguishable from professional content, we are not interested in all UGC content, but only in the quality that clearly differs from the professional. For this content, we use the term “in the wild” as a concept closely related to the concept of UGC, which is its special case. In this paper, we show that it is possible to deliver the new concept of an objective “in-the-wild” video content recognition model. The value of the F measure in our model is 0.988. The resulting model is trained and tested with the use of video sequence databases containing professional and “in the wild” content. These modelling results are obtained when the random forest learning method is used. However, it should be noted that the use of the more explainable decision tree learning method does not cause a significant decrease in the value of measure F (an F-measure of 0.973).
The growth in video streaming has been an exponential one for the last decade or so. High-resolution videos require high bandwidth to transport the videos over the network. There has been a growing demand for compression technologies to compress videos while simultaneously maintaining quality. Video codecs are used to encode and decode video streams. These codecs have been developed byMPEG, Google, Microsoft, andApple Inc. There aremany encoding parameters that affect bitrate and video quality. These performance parameters must be exploited, evaluated, and modeled to find the best possible solutions. This paper demonstrates some preliminary results for video coding sets with selected bitrates. The objective video multimethod assessment fusion (VMAF) metric is calculated for the encoded video versions. In this study, the quality of the encoded videos was evaluated and estimated using VMAF. The results confirm a strong relationship between bitrate and VMAF estimates. This study shows the impact of coding parameters on the VMAF values and provides the foundation for building robust models in the field of video quality analysis.
To evaluate a system that automatically summarizes video files (image and audio) and text, how the system works, and the quality of the results should be considered. With this objective, the authors have performed two types of evaluation: objective and subjective. The actual assessment is performed mainly automatically, while the individual assessment is based directly on the opinion of people, who evaluate the system by answering a set of questions, which are then processed to obtain the targeted conclusions. One of the purposes of the described research is to try to narrow the space of possible summarization scenarios. Meanwhile, in the light of individual results obtained, the researchers cannot unambiguously indicate one single scenario, recommended as the only one for further development. However, the researchers can state with certainty that the new development of scene 1, which has received many negative evaluations among professionals, should be discontinued. Considering the results of the set of questions about the quality of the complete system, the end-users have evaluated the scenario 3, and they think that the quality is excellent, obtaining results over 70
In this paper we present our work towards an effective solution for detection of dangerous objects, such as firearms or knives in a Closed Circuit Television System. We have gathered a large, manually annotated dataset of recordings supplemented by our original artificial sample generation method. We have used this dataset for training of a convolutional neural network. We present our approach and training results. We have also implemented and present software architecture that implements the neural network. We have shown, that the convolutional neural networks are well suited even for such complex object detection task, when provided with enough training samples.
The aim of the work is to report the results of the Chist-Era project AMIS (Access Multilingual Information opinionS). The purpose of AMIS is to answer the following question: How to make the information in a foreign language accessible for everyone? This issue is not limited to translate a source video into a target language video since the objective is to provide only the main idea of an Arabic video in English. This objective necessitates developing research in several areas that are not, all arrived at a maturity state: Video summarization, Speech recognition, Machine translation, Audio summarization and Speech segmentation. In this article we present several possible architectures to achieve our objective, yet we focus on only one of them. The scientific locks are be presented, and we explain how to deal with them. One of the big challenges of this work is to conceive a way to evaluate objectively a system composed of several components knowing that each of them has its limits and can propagate errors through the first component. Also, a subjective evaluation procedure is proposed in which several annotators have been mobilized to test the quality of the achieved summaries.
In this paper we present the results of the integration works on the system designed for automated summarization and translation of newscast and reports. We show the proposed system architectures and list the available software modules. Thanks to well defined interfaces the software modules may be used as building blocks allowing easy experimentation with different summarization scenarios.
In this paper, we present the first results of the project AMIS (Access Multilingual Information opinionS) funded by Chist-Era. The main goal of this project is to understand the content of a video in a foreign language. In this work, we consider the understanding process, such as the aptitude to capture the most important ideas contained in a media expressed in a foreign language. In other words, the understanding will be approached by the global meaning of the content of a support and not by the meaning of each fragment of a video. Several stumbling points remain before reaching the fixed goal. They concern the following aspects: Video summarization, Speech recognition, Machine translation and Speech segmentation. All these issues will be discussed and the methods used to develop each of these components will be presented. A first implementation is achieved and each component of this system is evaluated on a representative test data. We propose also a protocol for a global subjective evaluation of AMIS.
This paper presents a framework for summarization for newscasts and reports, that is a part of an ongoing research towards multilingual opinion analysis system conducted under the CHIST-ERA project “Access Multilingual Information opinionS” (AMIS). We present the results of qualitative analysis of newscast and reports published in the Internet by leading English, French and Arabic TV channels. We show the method used for creation of a database that contains 300 h of such content on controversial topics. Finally, we show the design and operation of our summarization framework. The framework is designed in such a way, that it allows for easy experimentation with different approaches to video summarization. The description is followed by a presentation of high- and low-level metadata extraction algorithms that include detection of the anchorperson, recognition of day and night shots and extraction of low-level video quality indicators.
The paper presents a new approach to the problem of calibration of a pair of CCTV cameras – a wide angle camera and Pan-Tilt-Zoom (PTZ) camera. The proposed solution allows for accurate control of the PTZ camera using point of interest selected in the coordinate system of the wide angle camera. In the paper we describe the local feature based calibration algorithm, camera control algorithm. We present results of the tests conducted with a real-life a setup of two high end CCTV cameras overseeing a 3000 m2 parking lot. We have achieved an average 1,24 ^∘ accuracy of calibration that translates to approximately 1.7 m at 80 m of observation distance. The proposed solution is designed in such a way, that after a one-time, heavy-computing calibration the camera control procedure is instantaneous and thus very well suited for operation with advanced object detection algorithms.
The paper presents application of multinomial logistic regression for color segmentation. The common problem in the subject of image understanding is creation of a large enough corpus for algorithm training. Especially when a large set of classes has to be recognized or if using convolutional neural networks the size and diversity of the training set strongly influences the quality of the resulting system. We present a method of automated generation of training samples by combining a well-known green box technique with multinomial logistic regression for background substitution. We show the encountered problems and their solutions. We present numerous examples of algorithm performance in background substitution. We conclude the paper with presentation of other examples of application of logistic regression for image understanding.
In this paper, we will present a study concerning the understanding of the needs of people using Internet in order to access to multilingual information. In fact, in the framework of AMIS (Accessing Multilingual Information and opinionS), a Chist-Era project, we propose to develop a system which will help to understand the main idea of a video in a foreign language. In order to design a useful system, a survey allowing to specify the profile of potential users of AMIS has been conducted. The study concerned 170 people from different countries: Poland, Spain and France. The sample is composed of people of different ages and different culture and languages.The results, in terms of requirements, achieved from this study show differences depending on how often the people watch the news on TV or review them on the Internet, and on the age of the target group. These concrete results help us in several decisions concerning how to build a realistic architecture of AMIS.
Kamel Smaili合作论文数University - Computer Science Department6
Fernando Boavida合作论文数University of Coimbra2