The interactive nature of virtual reality (VR) challenges many assumptions of conventional audiovisual quality evaluation approaches, necessitating tools and methods that account for user agency, temporal coupling, and context. We present the Quality and Experience Evaluation (QExE) Tool for interactive VR. Several quality evaluation methods, additional questionnaires, behavioral and interactivity data collection, and the ability to load multiple suitable audio rendering plug-ins are included. The tool streamlines the evaluation process by automatically creating test items that include audiovisual media such as object-based or multichannel audio and three-dimensional (3D) spatial scenes or 360° videos. Utilizing a 3D game engine in tandem with control software, the method, questionnaire, and interactivity data can be saved to subject-specific sub-directories. Prior research utilizing the QExE tool is described, along with a novel case study investigating the effects of a VR training scheme on cognitive load, audio plausibility, and interaction behavior. This study demonstrates that the QExE tool can be effectively employed to collect perceptual, cognitive, and behavioral data for subjective evaluations in interactive VR environments.
Cross-laboratory subjective quality studies rely on Mean Opinion Scores (MOS) as ground truth, yet systematic differences between laboratories persist even after alignment. The VQEG Multimedia Phase 2 study showed that MOS is relative rather than absolute, with individual subject differences being the second-most important source of variability - but these were never quantified. We use SUREAL to recover subject-level bias parameters and decompose variance through pooled vs. perlaboratory analyses. Results show that 85% of bias variance reflects genuine subject differences, while 15% is attributable to lab-level systematic effects. We also find that content-dependent laboratory effects are strongest for animated and music content. The substantial within-country variation suggests that laboratory-specific factors may mask cultural influences rather than negate them. These findings argue for reporting subject-level parameter distributions alongside MOS.
Recent video super-resolution (VSR) approaches use deep neural networks to enhance low-quality input videos and recover visual detail, with diffusion-based methods in particular showing promising results. In this paper, we investigate whether existing video quality models can be used to assess the performance of these diffusion-based VSR methods, by comparing model predictions with results from a subjective test. The study compares six upscaling methods (Lanczos, Rhea, SCST, DOVE, SeedVR2, Starlight Mini) applied to both compressed (AV1 and DCVC-RT) and uncompressed low-resolution videos considering the play-out on a UHD-1/4K screen. A range of full- and no-reference quality models are used to assess their applicability to this new type of quality degradation, focusing on within-sequence performance. The results highlight that CNN-based full-reference models, such as LPIPS, DISTS, and CVQA-FR show significantly higher correlation coefficients than both conventional full- as well as the tested no-reference models. Most overestimate the overly sharp results of SCST, with VMAF mainly failing due to spatial inconsistencies introduced by Starlight Mini. None of the tested video quality models reach sufficient accuracy so as to replace complementary subjective testing. The reference, degraded and upscaled videos, as well as the user ratings and model scores are made available with the paper at https://github.com/Telecommunication-Telemedia-Assessment/AVT-VQDB-UHD-1-VSR as open data.
This study investigated the influence of room acoustics, specifically reverberation time, on switching auditory selective attention. Two classroom scenarios were simulated using the room acoustic software raven: one with favorable acoustic properties (T30=0.48 s) and one with unfavorable properties (T30=1.52 s). In a listening experiment with 16 adult participants, both scenarios were live-rendered using the ravenintegration for Virtual Acoustics in Unity to examine auditory attention switching in an audiovisual virtual classroom. Reaction times were unaffected by reverberation. However, participants made significantly more errors in the unfavorable condition compared to the favorable and anechoic scenarios. Results further suggest that long reverberation times impaired the spatial separation between target and distractor sounds, leading to reduced performance in the attention-switching task. Post-experiment questionnaires indicated lower perceived speech intelligibility and higher listening effort in the highly reverberant condition. These findings demonstrate that reverberation consistent with realistic classroom acoustics can impair auditory selective attention in an audiovisual virtual reality scenario, highlighting the importance of incorporating room acoustic conditions when investigating complex listening tasks.
Traditionally, video quality models used in applications such as real-time service monitoring in video streaming or for encoder optimization are developed under specific practical constraints. One such constraint is the use of optimal viewing distances for specific target resolutions in subjective tests later acting as ground truth for developing quality models. This raises the question of how viewing distances different from the optimal affect the perceived video quality and hence the performance of state-of-the-art image and video quality models. To address this question, we present two subjective video quality tests (Study 1 and Study 2) for the 4K / UHD-1 screen resolution. Both studies used four viewing distances ranging from 1.6H to 4.8 H, with H being the screen height.The stimuli were encoded with H.265 at different bitrates ranging between 100 to 70000 kbit/s and four different resolutions from 360p to 2160p. We introduce our new 4K HDR database AVT-VQDB-Faces that is used in Study 2 in addition to the AVT-VQDB-UHD-1-VD that includes all test data. Study 1 (28 subjects) employed a constant bitrate encoding paradigm, while Study 2 (40 subjects) used a constant rate factor paradigm. While video quality typically improves with distance at moderate bitrates (400 to 3500 kbit/s), this trend reverses at high bitrates, proving that viewing distance determines the threshold where additional encoding precision no longer yields perceptible benefits. Thirteen image and video quality models, including both bitstream and pixel-based approaches, were evaluated for the impact of viewing distance on prediction accuracy. Results indicated that ITU-T Rec. P.1204.3 performed best overall across all distances, while VMAF and ADM2 performed best among pixel models in terms of Pearson Correlation and RMSE. Notably, model performance declined with an increased viewing distance, highlighting the need to incorporate viewing distance into video quality prediction. This is especially important since subjects typically chose an average preferred viewing distance of 3.4 H.
Telepresence robots that allow communication between older adults and their remotely located social contacts can foster social integration. The present laboratory test study explores older adults’ successful use of a telepresence robot (Research Question 1 [RQ1]), as well as their perceived enjoyment (RQ2), perceived ease of use (RQ3), perceived usefulness (RQ4), perceived social presence (RQ5), and intention to use (RQ6) a telepresence robot for robot-mediated communication (RMC). Semi-structured interviews, observations, and questionnaires were applied with a group of N = 14 older adults living in Germany. Participants completed a navigational task (as remote users) and an interpersonal communication task (as local users). Results show older adults used the telepresence robot successfully (RQ1) during the tasks. Furthermore, in interviews, older adults described their perceived enjoyment (RQ2), perceived ease of use (RQ3), and perceived usefulness (RQ4) during RMC as generally high. Perceived social presence (RQ5) during RMC was generally described as high, with RMC being considered a viable substitute when face-to-face communication is not possible. Finally, only two participants (2/14) had no intention to use (RQ6) a telepresence robot in the long term. Future design recommendations are provided, such as adapting the telepresence robot’s functions to older adults physical, psychological, and social conditions.
Emerging communication systems explore spatially enriched representations, e.g., including 3D representations, as extensions or replacements for conventional 2D video conferencing. In this paper, we evaluated a real-time floating-head communication framework comprising three modalities: classical video communication (VC-classic), a 2D projected floating-head representation (FH-2D), and a 3D floating-head variant (FH-3D). The system reconstructs a lightweight 3D facial mesh from monocular camera input using machine-learning-based facial feature extraction and real-time rendering for two participants in a standard teleconferencing setting using network connections. FH-3D introduces depth cues and is shown with 3D displays in contrast to FH-2D, where only 2D projected variants of the heads are shown. To evaluate the new representations in comparison to VC-classic, we conducted a user study employing a collaborative task scenario. Participants rated perceived presence, engagement, communication quality, and comfort. The results indicate that e.g., FH-3D improves spatial presence-related measures compared to FH-2D and enhances telepresence, while VC-classic remains strong in terms of familiarity, overall communication comfort, and is still preferred in comparison to head-only representations. Our findings contribute to the understanding of design choices and user experience in real-time camera-based telepresence systems.
Recent technological advances have increased the popularity of immersive video consumption, such as headmounted devices for 360 degrees video displays. Higher video resolutions and highly realistic scene representations are typically prioritized to enhance immersion in 360 degrees videos. One often not-so-wellstudied and analyzed aspect related to the sense of immersion in 360 degrees videos is the video stream paired with spatial audio, which is addressed in this paper. Two subjective studies are designed and conducted to investigate the interplay between audio and encoding parameters on video quality perception and users' exploration behavior. The studies follow a within-subjects and between-subjects design to assess the impact of 4(th)-order Ambisonics audio (4OA audio) on the perceived 360 degrees video quality, exploration behavior, and user's sense of comfort measured as cybersickness. In the experiments conducted, audio representation (4OA audio and no audio) and video compression parameters (H.264 and H.265 codecs with different settings) were varied. The audio conditions were found not to influence the video quality rating and the cybersickness in the test carried out. This finding demonstrates that participants were able to focus on visual quality during the assessment task without being influenced by the presence of spatial audio. At the same time, increased scene exploration with 4OA audio suggests that spatial audio did affect how participants interacted with the content, highlighting its role in shaping engagement even when not directly reflected in quality ratings. Additionally, less exploration was observed during repeated viewings, indicating that participants became familiar with the content and adopted a more focused viewing strategy to judge visual quality.
Group-to-group telepresence systems immerse geographically separated groups in a shared interaction space where remote users are represented as avatars. Notably, such systems allow users to interact with collocated and remote interlocutors simultaneously. In this context, where virtual user representations can be directly compared with real users, we investigate how visual realism (avatar type) and aural realism (presence of spatial audio) affect communication. Furthermore, we examine how communication differs between collocated and remote pairs of interlocutors. In our user study, groups of four participants perform a collaborative conversation task under the aforementioned visual and aural realism conditions. Our results indicate that avatar realism has positive effects on subjective ratings of perceived message understanding and group cohesion, and yields behavioural differences that indicate more interactivity and engagement. Few significant effects of aural realism were observed. Comparisons between collocated and remote communication found that collocated communication was perceived as more effective, but that more visual attention was paid to both remote participants than the collocated user.
We present ICS-MR, a dataset containing three conversational scenarios designed for the evaluation of communication quality in Mixed Reality (MR) systems. Along with detailed descriptions of the conversation tasks, we provide all the materials required to incorporate the tasks into MR user studies. The materials also support application of the scenarios in real-world and video-conferencing contexts for studies that, for example, call for comparison of immersive systems against reference communication media. Open-source Unity implementations of the scenarios are also made available, supporting direct usage of the scenarios in distributed, multi-user experiments. The conversation tasks have all been administered in recent scientific works that address the evaluation of user experiences in immersive communication systems, allowing analysis and comparison of each scenario's evoked behavioral properties. The ICS-MR dataset therefore contributes valuable resources for further research on communication in immersive systems.
Subjective testing is widely used for visual quality assessment to evaluate the impact of both technical and non-technical factors on user perception. Although standardized methods exist for collecting subjective ratings of visual quality, these ratings are inevitably influenced by each subject’s accuracy, manifesting as subject bias and inconsistency. Recommendations such as ITU-T P.910 and ITU-R BT.500 propose standardized methods to remove bias from subjective ratings. For instance, Annex E of ITU-T Rec. P.910 provides an effective strategy for addressing both subject bias and inconsistency. In this paper, we analyze 29 different visual quality assessment studies conducted over an eight-year period to understand the rating behaviour of subjects using these methods. Our investigation focuses on subjective studies targeting 4K, 8K, as well as 360° video and high-resolution image quality assessment. In the context of 4K video quality assessment, both SDR and HDR evaluations are considered. Our dataset covers use cases of short-term video quality and overall session quality assessment for HTTP-based adaptive streaming (HAS). For both these use cases, we propose a range of subject bias and inconsistency values that can serve as a reference for future studies. Furthermore, we define six different measures that can be used to assess the reliability of future studies and provide reference values for these measures. Following an open-science approach, all individual ratings from the included subjective tests, along with the results of the large-scale analysis, are made publicly available with this paper.
Introduction Omnidirectional videos (ODV) have become increasingly relevant in multimedia research and applications. Most prior studies focused on viewing these videos with head-mounted displays (HMDs) in controlled lab environments, requiring complex setups and costly hardware. This limits the feasibility of large-scale ODV studies for user behaviour analysis, as not everyone owns an HMD at home. Since ODV salience reflects exploration behaviour more closely aligned with real-world viewing compared to 2D video, alternative and more scalable approaches are of high interest. Methodology Towards this end, this paper analyses to what extent exploration behaviour observed in screen-based setups (lab, out-of-the-lab, and crowdsourcing) is similar to HMD-based viewing. A framework for screen-based ODV playback and data collection was developed, using mouse movements to change the corresponding viewport. Using 20 ODV sequences, three screen-based tests were conducted: a crowd-sourced test with 45 participants, an out-of-the-lab test with 45 participants, and a lab test with 33 participants. The results were compared with an existing HMD-based lab study using a similar design. Results and conclusions For evaluation of the user behaviour, Similarity Ring Metric (SRM) scores and intentional head movement counts are used. The findings show that exploration patterns in screen-based viewing differ from those in HMD viewing, meaning the two approaches are not fully equivalent. Nevertheless, for specific applications such as general viewport prediction, screen-based data can provide a viable alternative, offering potential for improved viewport recommendations in non-immersive viewing scenarios. Furthermore, using the fixation maps from the behavioural data of all tests six saliency models have been evaluated using established evaluation metrics such as the correlation coefficient. In the HMD-based lab test, SalViT360 achieved the highest prediction accuracy (especially at higher frame aggregation levels), followed by SUM, MDS-ViTNet, and DeepGaze IIE, while UNISAL performed worst but still outperformed V-BMS360, whereas all models showed reduced accuracy in the screen-based tests. To foster reproducibility and open science, both the collected datasets and the adapted framework for screen-based testing have been made publicly available.
Staying socially connected is vital for older adults' wellbeing. This study explores how wearable augmented reality (AR) can enhance their communication experiences. N=16 older adults from Germany used a prototype wearable AR communication system to engage in a conversation task with a realistic avatar representing a remote communication partner. Their experiences, focusing on observed user behavior, perceived usability, user engagement, social presence, and intention to use, were assessed through questionnaires, scales, think-aloud protocols, and interviews. Overall, participants reported positive and engaging experiences. Despite concerns about the headset's ergonomic fit and the avatar's limited emotional expressiveness, a high degree of social presence was achieved. Most participants were willing to use the system, particularly if social isolation became a concern. Our findings demonstrate wearable AR's potential to improve interpersonal communication among older adults and provide design insights to better address their social and usability needs. Specifically, it emphasizes the importance of enhancing the avatar's emotional expressiveness, particularly when interacting with familiar individuals, to effectively engage older adults and foster communication satisfaction.
Omnidirectional videos, or 360(circle) -videos represent an entry point to immersive media. However, watching 360(circle) -videos in head-mounted displays may evoke symptoms related to cybersickness. This literature review investigates the underexplored area of cybersickness induced by 360(circle) videos when viewed through head-mounted displays. While cybersickness has been studied extensively in interactive virtual reality (6 DOF), its manifestations and triggers in the area of 360(circle) -videos (3 DOF) remain less understood. By evaluating 52 selected articles, we discerned and categorized key factors influencing cybersickness into four groups: Technical, human-related, context and content, and experience-related factors. Overall, our review highlights that the specific characteristics of 360(circle) -video, such as recording and streaming technologies, limited agency, and having only 3-DOF, have been neglected in cybersickness research so far. By that, this review emphasizes the need for more systematic (longterm) studies about influencing factors. We also highlight the need for models predicting and estimating cybersickness and longitudinal studies to enhance the 360(circle) -video experience and ensure its broader acceptance beyond academia.
Besides traditional 2D video communication approaches, new systems aim to create more realistic and immersive representations of the involved conversation partners. This work presents a real-time floating-head communication setup that enables spatially separated visualizations of remote participants using only a standard camera and display. The underlying reconstruction pipeline applies machine-learning-assisted facial feature extraction to infer a 3D mesh of the participant's head from a live 2D video stream. Texturing maintains visual fidelity while supporting real-time performance. A dedicated transmission pipeline enables the exchange of 3D and texture data over conventional network connections, allowing flexible and location-independent use. A first lab test with eight participant pairs was performed to evaluate the system during a collaborative communication task. Subjective assessments using established telepresence and quality questionnaires confirmed the technical feasibility of the approach and its potential to enhance the sense of spatial presence. However, the overall perceptual quality and comfort did not yet reach the level of classical 2 D video communication. The study demonstrates the promise of accessible, spatially expressive communication setups that may bridge the gap between conventional video calls and emerging volumetric telepresence systems.
Interactive communication (IC), i.e., the reciprocal exchange of information between two or more interactive partners, is a fundamental part of human nature. As such, it has been studied across multiple scientific disciplines with different goals and methods. This article provides a cross-disciplinary and selective primer on contemporary IC integrating psychological mechanisms with speech signal, acoustic and media-technological constraints in theory, measurement, and applications. First, we outline theoretical frameworks that account for verbal, nonverbal, and multimodal aspects of IC, including distinctions between face-to-face and computer-mediated communication. Second, we summarize key methodological approaches, including behavioral, cognitive, and experiential measures of communicative synchrony and acoustic signal quality. Third, we discuss selected applications, applications in which speech transmission, signal enhancement, mediated dialogue, and real-time coordination are central, namely assistive listening technologies, conversational agents, and social VR, alongside ethical considerations. Taken together, this primer highlights how human capacities and technical systems jointly shape IC, consolidating concepts, findings, and challenges that have often been discussed in separate lines of research.
When humans listen to popular music, they often feel ‘the groove’, which is a pleasurable urge to move along with music. Previous studies suggest that groove arises from the brain’s entrainment to rhythmical patterns of music in the delta-theta and beta bands, which enables the brain to predict temporal structures. However, this notion has only been tested using simplified or artificial acoustic stimuli and melodies, rather than naturalistic popular music. In this study, we combined electroencephalography (EEG) measurements with two measures of neural entrainment to test whether neural entrainment in the delta, theta, and beta bands predicts individuals’ groove ratings of musical and clapping stimuli, which served as a control. We presented nine pop music songs and twelve MIDI-based clapping rhythms, which participants rated for their groove perception on a 5-point scale. Our results show that music songs received very variable groove ratings while ratings of clapping stimuli showed less variability. Univariate analyses demonstrated that both music and clapping stimuli consistently entrained neural oscillations in the delta–theta bands, but not beta bands, as measured by intertrial phase coherence (ITPC) and stimulus-brain coherence (SBC). ITPC and SBC in the theta band weakly, but significantly correlated with groove ratings in fronto-central clusters. Decoding of groove ratings from multivariate spatiotemporal-spectral EEG patterns of entrainment with curve-fitting and machine-learning models predicted groove perception with high accuracy, up to 0.8 Pearson correlation, where combined patterns of inter-trial phase coherence appeared more informative than patterns of stimulus-brain coherence. Our results support the theory that the perception of groove is rooted in the brain’s ability to entrain to ongoing dynamic temporal structures of music in the theta band. They further show that multivariate decoding models can be leveraged to predict subjective groove online from neural entrainment with high accuracy and temporal resolution. Summary ### Competing Interest Statement The authors have declared no competing interest. Deutsche Forschungsgemeinschaft, https://ror.org/018mejw64, 438822823