The purpose of this study is to explore the effects of an affective recommendation system on the developmental trajectories of prospective teachers' emotional patterns, integrated with a Simulated Virtual Classroom (SVC) platform called SimlnClass. SVC exposes teachers to a range of student discourses in the form of unexpected stimuli. Fifteen prospective teachers participated in a study consisting of two practicum sessions in the SVC. Participants did not receive any affective recommendation after the first session but did receive it after the second session. Additional data were collected during both sessions in the SVC, including the physiological responses, such as electroencephalogram (EEG), galvanic skin response (GSR), and facial expressions. L metric and Lag sequential analysis were employed in determining teachers' transitional emotional patterns. The results showed that participants did not maintain disgust after receiving affective recommendations, although they maintained sadness. This result indicates that the given affective recommendation has an inherent effect on negative emotions that are felt less intensely. Different or longer-term interventions may be needed for more intense and long-lasting negative discrete emotions such as sadness. Also, participants transitioned to happiness and sadness instead of maintaining their neutral status after receiving an affective recommendation. This result demonstrates that affective recommendations encourage participants to use the cognitive reappraisal necessary for emotion regulation. When the participants' emotional patterns are examined on the basis of student discourse, the results are more complex and the emotional patterns differ according to the function of the discourse triggered by virtual students.
Deep learning has proven to be an important element of modern data processing technology, which has found its application in many areas such as multimodal sensor data processing and understanding, data generation and anomaly detection. While the use of deep learning is booming in many real-world tasks, the internal processes of how it draws results is still uncertain. Understanding the data processing pathways within a deep neural network is important for transparency and better resource utilisation. In this paper, a method utilising information theoretic measures is used to reveal the typical learning patterns of convolutional neural networks, which are commonly used for image processing tasks. For this purpose, training samples, true labels and estimated labels are considered to be random variables. The mutual information and conditional entropy between these variables are then studied using information theoretical measures. This paper shows that more convolutional layers in the network improve its learning and unnecessarily higher numbers of convolutional layers do not improve the learning any further. The number of convolutional layers that need to be added to a neural network to gain the desired learning level can be determined with the help of theoretic information quantities including entropy, inequality and mutual information among the inputs to the network. The kernel size of convolutional layers only affects the learning speed of the network. This study also shows that where the dropout layer is applied to has no significant effects on the learning of networks with a lower dropout rate, and it is better placed immediately after the last convolutional layer with higher dropout rates.
Quality of Experience (QoE) is becoming an important factor of User-Centred Design (UCD). The deployment of pure technical measures such as Quality of Service (QoS) parameters to assess the quality of multimedia applications is phasing out due to the failure of those methods to quantify true user satisfaction. Though significant research results and several deployments have occurred and been realized over the last few years, focusing on QoE-based multimedia technologies, several issues both of theoretical and practical importance remain open. Accordingly, the papers of this Special Issue are significant contribution samples within the general ecosystem highlighted above, ranging from QoE in the capture, processing and consumption of next-generation multimedia applications. In particular, a total of five excellent articles have been accepted, following a rigorous review process, which address many of the aforementioned challenges and beyond.
The electroencephalogram (EEG) has great attraction in emotion recognition studies due to its resistance to deceptive actions of humans. This is one of the most significant advantages of brain signals in comparison to visual or speech signals in the emotion recognition context. A major challenge in EEG-based emotion recognition is that EEG recordings exhibit varying distributions for different people as well as for the same person at different time instances. This nonstationary nature of EEG limits the accuracy of it when subject independency is the priority. The aim of this study is to increase the subject-independent recognition accuracy by exploiting pretrained state-of-the-art Convolutional Neural Network (CNN) architectures. Unlike similar studies that extract spectral band power features from the EEG readings, raw EEG data is used in our study after applying windowing, pre-adjustments and normalization. Removing manual feature extraction from the training system overcomes the risk of eliminating hidden features in the raw data and helps leverage the deep neural network's power in uncovering unknown features. To improve the classification accuracy further, a median filter is used to eliminate the false detections along a prediction interval of emotions. This method yields a mean cross-subject accuracy of 86.56% and 78.34% on the Shanghai Jiao Tong University Emotion EEG Dataset (SEED) for two and three emotion classes, respectively. It also yields a mean cross-subject accuracy of 72.81% on the Database for Emotion Analysis using Physiological Signals (DEAP) and 81.8% on the Loughborough University Multimodal Emotion Dataset (LUMED) for two emotion classes. Furthermore, the recognition model that has been trained using the SEED dataset was tested with the DEAP dataset, which yields a mean prediction accuracy of 58.1% across all subjects and emotion classes. Results show that in terms of classification accuracy, the proposed approach is superior to, or on par with, the reference subject-independent EEG emotion recognition studies identified in literature and has limited complexity due to the elimination of the need for feature extraction.
Loughborough University Multimodal Emotion Database-2 (LUMED-2) is a new multimodal emotion dataset that was created by the researchers of Loughborough University, UK, and Hacettepe University, Turkey, by collecting simultaneous multimodal data from 13 participants (6 females and 7 males) by showing audio-visual stimuli. The total duration of all stimuli is 8 minutes and 50 seconds, which consist of short video clips selected from the web to elicit specific emotions. Between each video clip, in order to let the participants have a rest, a 20-second grey screen was showed. Although it is anticipated that each video clip elicits a distinct emotional state and thus can determine the label of the resulting emotion, in reality the same content might trigger differing emotions for different participants. Therefore, after each session, the participants were additionally asked to label the clips with the felt emotional state while watching them. Three different emotions were resulted from labelling: "sad", "neutral" and "happy". The facial expressions of the participants were captured using a webcam at a resolution of 640x480 and at 30 fps. Participants’ EEG data was captured using an ENOBIO 8-channels wireless EEG device, which has a temporal resolution of 500 Hz. We filtered EEG data for the frequency range [0, 75Hz] and applied baseline subtraction for each window. As for the peripheral physiological data, an EMPATICA E4 Wristband, powered by Bluetooth, was used to record participants’ GSR.This multimodal emotion database is produced and should be used for research purposes only.For any questions related to the data contained in this database, contact Dr Yucel Cimtay (yucel.cimtay@gmail.com)
Multimodal emotion recognition has gained traction in affective computing research community to overcome the limitations posed by the processing a single form of data and to increase recognition robustness. In this study, a novel emotion recognition system is introduced, which is based on multiple modalities including facial expressions, galvanic skin response (GSR) and electroencephalogram (EEG). This method follows a hybrid fusion strategy and yields a maximum one-subject-out accuracy of 81.2% and a mean accuracy of 74.2% on our bespoke multimodal emotion dataset (LUMED-2) for 3 emotion classes: sad, neutral and happy. Similarly, our approach yields a maximum one-subject-out accuracy of 91.5% and a mean accuracy of 53.8% on the Database for Emotion Analysis using Physiological Signals (DEAP) for varying numbers of emotion classes, 4 in average, including angry, disgust, afraid, happy, neutral, sad and surprised. The presented model is particularly useful in determining the correct emotional state in the case of natural deceptive facial expressions. In terms of emotion recognition accuracy, this study is superior to, or on par with, the reference subject-independent multimodal emotion recognition studies introduced in the literature.
Multi-view plus-depth-map (MVD) video streaming with autostereoscopic displays provides multi-user immersive media experiences. In this context, delivery of MVD representation to multiple clients remains a challenging problem because of the high-volume of data involved and the inherent limitations imposed by the delivery networks. To this end, this paper investigates the side information (SI) assisted adaptation algorithm using peer-to-peer (P2P) systems. P2P delivery systems for MVD video can maximize link utilization, preventing the transport of multiple video copies of the same packet for many users. However, the quality of experience (QoE) can be significantly degraded by dynamic variations caused by network congestions. To this end, our solution comprises the extraction of low-overhead metadata at the encoding server that is distributed through the P2P network as SI and used by P2P clients performing network adaptation. In the proposed adaptation strategy, pre-selected views are discarded at times of network congestion and reconstructed with an optimal reconstruction performance using the delivered SI and the delivered neighboring camera views. The experimental results show that the robustness of P2P multi-view streaming using the proposed adaptation scheme is significantly increased in the P2P network.
In this paper an error concealment (EC)-aware encoding scheme is proposed to improve the quality of decoded video in broadcast environments subject to transmission errors and data loss. The proposed scheme is based on a scalable coding approach where the best EC methods to be used at the decoder are optimally determined at the encoder and signalled to the decoder through supplemental enhancement information messages. Such optimal EC modes are found by simulating transmission losses, followed by a Lagrangian optimization of the signalling rate-EC distortion cost. A generalized saliency-weighted distortion is used and the residue between coded frames and their EC substitutes is encoded using a rate-controlled enhancement layer. When data loss occurs the signalling information is used by the decoder to improve the reconstruction quality. The simulation results show that the proposed method achieves consistent quality gains in comparison with other reference methods and previous works. Using only the EC mode signalling, i.e., without any residue transmitted in the enhancement layer, an average PSNR gain up to 2.95 dB is achieved, while using the full EC-aware scheme, i.e., including residue encoded in the enhancement layer, the proposed scheme outperforms other comparable methods, with a PSNR gain up to 3.79 dB.
One of the key performance targets on the European Commission's Digital Agenda is to provide at least 30-Mbit/s broadband coverage to all European households by 2020. The deployment of existing terrestrial technologies will not be able to satisfy the requirements in the most difficult-to-serve locations, either due to a lack of coverage in areas where the revenue potential for terrestrial service providers is too low or due to technological limitations that diminish the available throughput in rural environments. In this paper, we investigate a hybrid broadband system combining satellite and terrestrial access networks. The system design and the key building blocks of the intelligent routing entities (referred to as intelligent gateways) are presented. To justify the hybrid broadband system's performance subjectively, lab trials have been performed with an integrated multiple access network emulator and a variety of typical multimedia applications that have varying requirements. The results of the lab trials suggest that the quality of experience is consistently improved thanks to the utilisation of intelligent gateway devices, when compared with using a single access network at a time.
When it comes to evaluating perceptual quality of digital media for overall quality of experience assessment in immersive video applications, typically two main approaches stand out: Subjective and objective quality evaluation. On one hand, subjective quality evaluation offers the best representation of perceived video quality assessed by the real viewers. On the other hand, it consumes a significant amount of time and effort, due to the involvement of real users with lengthy and laborious assessment procedures. Thus, it is essential that an objective quality evaluation model is developed. The speed-up advantage offered by an objective quality evaluation model, which can predict the quality of rendered virtual views based on the depth maps used in the rendering process, allows for faster quality assessments for immersive video applications. This is particularly important given the lack of a suitable reference or ground truth for comparing the available depth maps, especially when live content services are offered in those applications. This paper presents a no-reference depth map quality evaluation model based on a proposed depth map edge confidence measurement technique to assist with accurately estimating the quality of rendered (virtual) views in immersive multi-view video content. The model is applied for depth image-based rendering in multi-view video format, providing comparable evaluation results to those existing in the literature, and often exceeding their performance.
The increased compression ratios achieved by the High Efficiency Video Coding (HEVC) standard lead to reduced robustness of coded streams, with increased susceptibility to network errors and consequent video quality degradation. This paper proposes a method based on a two-stage approach to improve the error robustness of HEVC streaming, by reducing temporal error propagation in the case of frame loss. The prediction mismatch that occurs at the decoder after frame loss is reduced through the following two stages. First, at the encoding stage, the reference pictures are dynamically selected based on constraining conditions and Lagrangian optimization, which distributes the use of reference pictures, by reducing the number of prediction units that depend on a single reference. Second, at the streaming stage, a motion vector (MV) prioritization algorithm, based on spatial dependencies, selects an optimal subset of MVs to be transmitted, redundantly, as side information to reduce mismatched MV predictions at the decoder. The simulation results show that the proposed method significantly reduces the effect of temporal error propagation. Compared with the reference HEVC, the proposed reference picture selection method is able to improve the video quality at low-packet-loss rates (e.g., 1%) using the same bitrate, achieving quality gains up to 2.3 dB for 10% of packet loss ratio. It is shown, for instance, that the redundant MVs are able to boost the performance achieving quality gains of 3 dB when compared with the reference HEVC, at the cost using 4% increase in total bitrate.
Due to the complexity of the natural world, a programmer cannot foresee all possible situations, a connected and autonomous vehicle (CAV) will face during its operation, and hence, CAVs will need to learn to make decisions autonomously. Due to the sensing of its surroundings and information exchanged with other vehicles and road infrastructure, a CAV will have access to large amounts of useful data. While different control algorithms have been proposed for CAVs, the benefits brought about by connectedness of autonomous vehicles to other vehicles and to the infrastructure, and its implications on policy learning has not been investigated in literature. This paper investigates a data driven driving policy learning framework through an agent-based modelling approaches. The contributions of the paper are two-fold. A dynamic programming framework is proposed for in-vehicle policy learning with and without connectivity to neighboring vehicles. The simulation results indicate that while a CAV can learn to make autonomous decisions, vehicle-to-vehicle (V2V) communication of information improves this capability. Furthermore, to overcome the limitations of sensing in a CAV, the paper proposes a novel concept for infrastructure-led policy learning and communication with autonomous vehicles. In infrastructure-led policy learning, road-side infrastructure senses and captures successful vehicle maneuvers and learns an optimal policy from those temporal sequences, and when a vehicle approaches the road-side unit, the policy is communicated to the CAV. Deep-imitation learning methodology is proposed to develop such an infrastructure-led policy learning framework.
In this paper1, we introduce the concept of Virtual Transcendence Experience (VTE) as a response to the interactions of several users sharing several immersive experiences through different media channels. For that, we review the current body of knowledge that has led to the development of a VTE system. This is followed by a discussion of current technical and design challenges that could support the implementation of this concept. This discussion has informed the VTE framework (VTEf), which integrates different layers of experiences, including the role of each user and the technical challenges involved. We conclude this paper with suggestions for two scenarios and recommendations for the implementation of a system that could support VTEs.
Multiview entertainment is the next step in 3D immersive media networking owing to its improved depth perception and free-viewpoint viewing capability whereby users can observe the scene from the desired viewpoint. This paper outlines a delivery system for multiview plus depth video, combining the broadcast and broadband networks. The digital video broadcast (DVB) network is used along with adaptive peer-to-peer (P2P) distribution over the Internet to deliver high-volume multimedia to users. The DVB network has been used to deliver part of the 3D service, owing to its robustness and wide availability, as a mechanism to guarantee the minimum 3D quality of experience. The developed system brings key contributions in the P2P transport for real-time multimedia delivery, including a user preference-aware adaptation mechanism, adaptive redundant chunk scheduling for robustness, incentives to decrease the load on the content server for improved system scalability, and resynchronization capability with the DVB transmission. The introduced features are compared with those of some other well-known P2P solutions to highlight the quantitative gains. A subjective testing campaign has also been organized on the developed hybrid platform, which proves the effectiveness of user-aware adaptation over network-based adaptation on a mean opinion score scale.
In this paper a robust encoding scheme is proposed to improve the visual quality of HEVC decoded video when intra frames are lost along the streaming path. For this purpose, the encoding process includes frame loss simulation and subsequent error concealment, to find the most efficient method that should be used by a decoder to recover lost intra frames. In this novel scheme, each image is divided into partitions, which are associated with the error concealment method that achieves the lowest distortion. Then this information is signalled to the decoder through SEI messages in the coded stream. In order to efficiently use the signalling overhead, rate-distortion optimisation is used to achieve the best trade-off between the number of transmitted symbols and distortion of reconstructed frames. Experimental results show the effectiveness of the proposed method to enhance the quality of reconstructed intra frames under different packet loss ratios (PLR). For PLR=10%, the robust coding scheme is able to improve the average PSNR of all frames affected by errors, up to 1.50 dB and 3.44 dB in Low-Delay and Random-Access configurations respectively, at a maximum overhead cost of 0.24%.
In this paper a fixation prediction based saliency algorithm is used in order to predict the head movements of viewers watching virtual reality (VR) videos, by modelling the relationship between fixation predictions and recorded head movements. The saliency algorithm is applied to viewings faithfully recreated from recorded head movements. Spherical cross-correlation analysis is performed between predicted attention centres and actual viewing centres in order to try and identify prevalent lengths of predictable attention and how early they can be predicted. The results show that fixation prediction based saliency analysis correlates with head movements only for limited durations. Therefore, further classification of durations where saliency analysis is predictive is required.
The increase in Internet bandwidth and the developments in 3D video technology have paved the way for the delivery of 3D Multi-View Video (MVV) over the Internet. However, large amounts of data and dynamic network conditions result in frequent network congestion, which may prevent video packets from being delivered on time. As a consequence, the 3D video experience may well be degraded unless content-aware precautionary mechanisms and adaptation methods are deployed. In this work, a novel adaptive MVV streaming method is introduced which addresses the future generation 3D immersive MVV experiences with multi-view displays. When the user experiences network congestion, making it necessary to perform adaptation, the rate-distortion optimum set of views that are pre-determined by the server, are truncated from the delivered MVV streams. In order to maintain high Quality of Experience (QoE) service during the frequent network congestion, the proposed method involves the calculation of low-overhead additional metadata that is delivered to the client. The proposed adaptive 3D MVV streaming solution is tested using the MPEG Dynamic Adaptive Streaming over HTTP (MPEG-DASH) standard. Both extensive objective and subjective evaluations are presented, showing that the proposed method provides significant quality enhancement under the adverse network conditions.
Advances in video coding and networking technologies have paved the way for the Multi-View Video (MVV) streaming. However, large amounts of data and dynamic network conditions result in frequent network congestion, which may prevent video packets from being delivered on time. As a consequence, the 3D viewing experience may be degraded significantly, unless quality-aware adaptation methods are deployed. There is no research work to discuss the MVV adaptation of decision strategy or provide a detailed analysis of a dynamic network environment. This work addresses the mentioned issues for MVV streaming over HTTP for emerging multi-view displays. In this research work, the effect of various adaptations of decision strategies are evaluated and, as a result, a new quality-aware adaptation method is designed. The proposed method is benefiting from layer based video coding in such a way that high Quality of Experience (QoE) is maintained in a cost-effective manner. The conducted experimental results on MVV streaming using the proposed strategy are showing that the perceptual 3D video quality, under adverse network conditions, is enhanced significantly as a result of the proposed quality-aware adaptation.
Vladan Velisavljevic合作论文数Deutsche Telekom Laboratories4