Immersive technologies promise to revolutionize communication through enhanced sense of presence and interactivity. To enable interaction, reliable low-latency transport mechanisms are needed to handle the large volumes of data created by complex 3D objects. In this paper, we propose an open-source, codec-independent, selective forwarding unit (SFU) for real-time volumetric video streaming using WebRTC. For evaluation purposes, we provide a reference client implementation by extending VR2Gather, a TCP-based system for immersive communication. We conduct extensive evaluations using both new and existing datasets to compare the performance of WebRTC against TCP-based protocols in an emulated testbed environment. The evaluations demonstrate that WebRTC outperforms other protocols in high-latency scenarios and adapts video quality to user movement 13% and 36% faster than its TCP-based counterparts in networks with 5 ms and 10 ms of network latency, respectively.
Recent advancements in immersive communication technologies have facilitated the exploration of novel modalities of social interaction, utilizing platforms that offer a spectrum of representations from simplified avatars to photorealistic volumetric and point-cloud reconstructions. This study focused on how these different representations and degrees of immersion impact communication processes and collaborative task performance. The primary objective was to examine potential variances in performance and interaction quality across different systems during a basic communication task. To achieve this, we employed Meta Horizon Workrooms as a three degrees of freedom (3DoF) system, VR2Gather as a six degrees of freedom (6DoF) system, and the MS Teams meeting application as a comparative benchmark. A user study was conducted whereby participants engaged in a charade game, alternating between the roles of “mimic” and “guesser.” Throughout the experimental session, data were collected to assess comfort, presence, task load, and interaction quality, with performance quantified by the number of words guessed per minute. The findings indicated no significant differences in workload, presence, or simulator sickness across platforms; however, significant differences were observed in audio–visual quality, communication effort, and enjoyment. Performance metrics revealed that participants exhibited the highest task performance within MS Teams, followed by VR2Gather, and subsequently Horizon Workrooms. These findings underscore the influence of embodiment fidelity and system familiarity on communication quality within immersive environments. We discuss system-level constraints and propose considerations for the design of future social extended reality (XR) platforms that effectively balance realism with usability and performance.
Social VR offers new possibilities for parent-child co-play through immersive and embodied interaction, but still little is known about how game design shapes collaboration in these environments. This study investigates how 1) game design impact parent-child interaction and collaboration in a social VR game – The Space Archivists – and what 2) the perceived added value is by parents and children experiencing social VR games together. We conducted closed and open tests of the game with ten parent-child pairs (N=20), and explored how they collaborate and interact in the virtual environment. Findings show that collaboration combined family interaction patterns with new forms of coordination introduced by immersive VR. Parents often organized tasks and provided guidance on how to complete them, while children frequently developed greater expertise with VR controls, resulting in flexible exchanges of leadership and expertise between parents and children. Participants used verbal- and embodied communication, for example describing the physical space to each other, pointing, avatar positioning, and proximity to establish shared attention and coordinate actions. However, difficulties in handling VR-controls, disengagement, and external assistance at times disrupted direct parent-child interaction. Overall, the findings suggest that the added value of social VR lies in creating opportunities for parents and children to collaborate, exchange complementary expertise, and connect through embodied communication. Social VR therefore extends joint media engagement by combining familiar family dynamics with new spatial and embodied forms of shared play.
Perceptual quality assessment of Dynamic Point Cloud (DPC) contents plays an important role in various Virtual Reality (VR) applications that involve human beings as the end user. Understanding and modeling perceptual quality assessment is greatly enriched by insights from visual attention. However, incorporating aspects of visual attention in DPC quality models is largely unexplored, as ground-truth visual attention data are scarcely available. Besides, testing methods and procedures for collecting visual attention data are still to be agreed on. This article presents a dataset containing subjective opinion scores and visual attention maps of DPCs, collected in a VR environment using eye-tracking technology. Both the quality score and eye-tracking data were collected during a subjective quality assessment experiment, in which subjects were instructed to watch and rate DPCs at various degradation levels under 6 Degrees of Freedom (DoF) inspection, using a head-mounted display. Qualitative interview analysis was also conducted after the experiment. The dataset consists of 50 DPCs, including 5 reference DPCs, with each reference encoded at 3 distortion levels using 3 different codecs (namely G-PCC, V-PCC, CWI-PCL), amounting to a total of 9 degraded version per reference. Additionally, it incorporates 1,000 gaze trials from 40 participants, yielding a total of 15,000 visual attention maps across all the DPCs. We additionally benchmark objective quality metrics originally designed for static point clouds, evaluating their performance in our dataset using two temporal pooling strategies. Furthermore, we employ the visual attention data that are retrieved during our experiment to evaluate whether the performance of widely used objective quality metrics is improved by considering subjective measurements of visual attention. This dataset establishes a link between quality assessment and visual attention within the context of DPC. Moreover, thematic analysis of the interviews helps uncover user behavior and factors impacting perceptual quality for DPC in 6 DoF. This work deepens our understanding of DPC quality assessment and visual attention, driving progress in the realm of VR experiences and perception.
In recent years, video conferencing platforms have become powerful tools for remote communication. There has also been an increase in the use of VR systems for communication. However, very few of these systems utilize photorealistic human representation. This paper investigates the strengths, challenges, and limitations of a novel 3D communication prototype (VR2Gather) and a well-established video conferencing system (Zoom). Specifically, we explore whether the 3D communication prototype can achieve comparable performance levels in a remote physiotherapy use case. By assessing various aspects, such as audio-visual quality, presence, and interaction, we aim to determine if the current prototype is comparable with commercial systems in some dimensions while exceeding expectations in others. Our results indicated that VR2Gather has the potential for a better sense of connection and higher concentration. However, challenges like improving 3D rendering quality and communication ease still need to be overcome to make it suitable for physiotherapy.
Social VR allows users to interact with each other and explore virtual spaces together. This demo presents a social VR fashion museum, where visitors engage with fashion artefacts. The experience spans across three different spaces designed for various goals. This application, created following a human-centred approach, explores how the visitors interact with each other and the exhibits, considering their 3D volumetric representation and changing environmental context throughout the experience.
Extended Reality (XR) systems are rapidly shifting from isolated, single-user applications towards collaborative and social multi-user experiences. To evaluate the quality and effectiveness of such interactions, it is therefore required to move beyond traditional individual metrics such as Quality-of-Experience (QoE) or Sense of Presence (SoP). Instead, group-level dynamics such as effective communication, coordination etc. need to be encompassed to assess the shared understanding of goals and procedures. In psychology, this is referred to as a Shared Mental Model (SMM). The strength and congruence of such an SMM are known to be key for effective team collaboration and performance. In an immersive XR setting, though, novel Influence Factors (IFs) emerge that are not considered in a setting of physical co-location. Evaluations on the impact of these novel factors on SMM formation in XR, however, are close to non-existent. Therefore, this work proposes SMMs as a novel evaluation tool for collaborative and social XR experiences. To better understand how to explore this construct, we ran a prototypical experiment based on ITU recommendations in which the influence of asymmetric end-to-end latency is evaluated through a collaborative, two-user block building task. The results show how also in an XR context strong SMM formation can take place even when collaborators have fundamentally different responsibilities and behavior. Moreover, the study confirms previous findings by showing in an XR context that a teams’ SMM strength is positively associated with its performance.
Volumetric video is a key enabler of immersive extended reality (XR) experiences and is often represented using point clouds for their structural simplicity. However, capturing volumetric content through multi-view acquisition and depth sensing poses many challenges, such as occlusions and depth mismatches. To foster research in this field, we introduce a unique dual-quality point cloud dataset, named UVG-CWI-DQPC, which is designed to support the development of point cloud enhancement, compression, and quality assessment. Our dataset includes 12 dynamic sequences captured simultaneously by: 1) a high-end capture system producing high-fidelity point clouds with extensive processing; and 2) a consumer-grade capture system relying on affordable RGB-D cameras, lightweight processing, and open-source tools. For each sequence, our dataset provides ground-truth point clouds from the high-end capture system and raw RGB-D footage from the consumer-grade capture system, along with calibration data and tools for point cloud generation. This dual-quality setup enables direct comparison and benchmarking of algorithms for densification, occlusion removal, registration, and quality enhancement. Our dataset is publicly available under a permissive license to support reproducible research and standardization work in Moving Picture Experts Group (MPEG) and 3rd Generation Partnership Project (3GPP).
Social Virtual Reality is envisioned to transform how individuals communicate remotely, offering a sense of immersion and co-presence within a virtual space. Current platforms enabling remote social interactions rely on synthetic user representations. We address this limitation by enabling realistic human representation through volumetric content capture, encoding and transmission. Specifically, we present an extended version of VR2Gather, now a fully open source Unity package, available at https://github.com/ cwi-dis/VR2Gather-acmmm-oss. Our platform is a customisable system to transmit volumetric content in a multi-party real-time environment, easy to integrate into existing applications.
Immersive technologies like eXtended Reality (XR) are the next step in videoconferencing. In this context, understanding the effect of delay on communication is crucial. This article presents the first study on the impact of delay on collaborative tasks using a realistic Social XR system. Specifically, we design an experiment and evaluate the impact of end-to-end delays of 300, 600, 900, 1,200, and 1,500 ms on the execution of a standardized task involving the collaboration of two remote users that meet in a virtual space and construct block-based shapes. To measure the impact of the delay in this communication scenario, objective and subjective data were collected. As objective data, we measured the time required to execute the tasks and computed conversational characteristics by analyzing the recorded audio signals. As subjective data, a questionnaire was prepared and completed by every user to evaluate different factors such as overall quality, perception of delay, annoyance using the system, level of presence, cybersickness, and other subjective factors associated with social interaction. The results show a clear influence of the delay on the perceived quality and a significant negative effect as the delay increases. Specifically, the results indicate that the acceptable threshold for end-to-end delay should not exceed 900 ms. This article additionally provides guidelines for developing standardized XR tasks for assessing interaction in Social XR environments.
Virtual reality telecommunication systems promise to overcome the limitations of current real-time teleconferencing solutions by enabling a better sense of immersion and fostering more natural interpersonal interactions. Many solutions that currently enable immersive teleconferencing employ synthetic avatars to represent their users. However, photorealistic reconstructions have been shown to increase the sense of presence with respect to synthetic avatars in teleimmersive scenarios. In this article, we present VR2Gather, a costumizable, end-to-end system to transmit volumetric contents in multiparty, real-time communication. We present the architecture and evaluate the costs and benefits of using different modules and transport mechanisms in terms of CPU usage, latency, and bandwidth. Moreover, we report the user experience based on applications the system has been used for and how it was customized to meet the requirements using different acquisition and rendering modules.
—Virtual Reality telecommunication systems promise to overcome the limitations of current real-time teleconferencing solutions, by enabling a better sense of immersion and fostering more natural interpersonal interactions. Many solutions that currently enable immersive teleconferencing employ synthetic avatars to represent their users. However, photorealistic reconstructions have been shown to increase the sense of presence with respect to synthetic avatars in tele-immersive scenarios. In this paper, we present VR2Gather, a costumizable end-to-end system to transmit volumetric contents in multi-party real-time communication. We present the architecture and evaluate the costs and benefits of using different modules and transport mechanisms in terms of CPU usage, latency, and bandwidth. Moreover, we report the user experience based on applications the system has been used for, and how it was customised to meet the requirements using different acquisition and rendering modules.
Perceptual quality assessment of Dynamic Point Cloud (DPC) contents plays an important role in various Virtual Reality (VR) applications that involve human beings as the end user, understanding and modeling perceptual quality assessment is greatly enriched by insights from visual attention. However, incorporating aspects of visual attention in DPC quality models is largely unexplored, as ground-truth visual attention data is scarcely available. This paper presents a dataset containing subjective opinion scores and visual attention maps of DPCs, collected in a VR environment using eye-tracking technology. The data was collected during a subjective quality assessment experiment, in which subjects were instructed to watch and rate DPCs at various degradation levels under 6 degrees-of-freedom inspection, using a head-mounted display. The dataset comprises 5 reference DPC contents, with each reference encoded at 3 distortion levels using 3 different codecs, amounting to a total of 9 degraded DPC contents. Moreover, it includes 1,000 gaze trials from 40 participants, resulting in 15,000 visual attention maps in total. The curated dataset can serve as authentic benchmark data for assessing the performance of objective DPC quality metrics. Additionally, it establishes a link between quality assessment and visual attention within the context of DPC. This work deepens our understanding of DPC quality and visual attention, driving progress in the realm of VR experiences and perception.
Social virtual reality (VR) allows multiple remote users to interact in a shared space, unveiling new possibilities for communication in immersive environments. Mediascape XR presents a social VR experience that teleports 3D representations of remote users, using volumetric video, to a virtual museum. It enables visitors to interact with cultural heritage artifacts while allowing social interactions in real time between them. The application is designed following a human-centered approach, enabling an interactive, educating, and entertaining experience.
Technological advances in head-mounted displays and novel real-time 3D acquisition and reconstruction solutions have fostered the development of 6 Degrees of Freedom (6DoF) teleimmersive systems for social VR applications. Point clouds have emerged as a popular format for such applications, owing to their simplicity and versatility; yet, dense point cloud contents are too large to deliver directly over bandwidth-limited networks. In this context, user-adaptive delivery mechanisms are a promising solution to exploit the increased range of motion offered by 6DoF VR applications to yield gains in perceived quality of 3D point cloud user representations, while reducing their bandwidth requirements. In this paper, we perform a user study in VR to quantify the gains adaptive tile selection strategies can bring with respect to non-adaptive solutions. In particular, we define an auxiliary utility function, we employ established methods from the literature and newly-proposed schemes for distributing the bit budget across the tiles, and we evaluate them together with non-adaptive streaming baselines through subjective QoE assessment. Results confirm that considerable gains can be obtained with user-adaptive streaming, achieving bit rate gains of up to 65% with respect to a non-adaptive approach to deliver comparable quality. Our analysis provides useful insights for the design and development of social VR applications.
Social virtual reality (VR) allows multiple remote users to interact in a shared space, unveiling new possibilities for communication in immersive environments. Mediascape XR presents a social VR experience that teleports 3D representations of remote users, using volumetric video, to a virtual museum. It enables visitors to interact with cultural heritage artifacts while allowing social interactions in real time between them. The application is designed following a human-centered approach, enabling an interactive, educating, and entertaining experience.
Remote communication has rapidly become a part of everyday life in both professional and personal contexts. However, popular video conferencing applications present limitations in terms of quality of communication, immersion and social meaning. VR remote communication applications offer a greater sense of co-presence and mutual sensing of emotions between remote users. Previous research on these applications has shown that realistic point cloud user reconstructions offer better immersion and communication as compared to synthetic user avatars. However, photorealistic point clouds require a large volume of data per frame and are challenging to transmit over bandwidth-limited networks. Recent research has demonstrated significant improvements to perceived quality by optimizing the usage of bandwidth based on the position and orientation of the user's viewport with user-adaptive streaming. In this work, we developed a real-time VR communication application with an adaptation engine that features tiled user-adaptive streaming based on user behaviour. The application also supports traditional network adaptive streaming. The contribution of this work is to evaluate the impact of tiled user-adaptive streaming on quality of communication, visual quality, system performance and task completion in a functional live VR remote communication system. We performed a subjective evaluation with 33 users to compare the different streaming conditions with a neck exercise training task. As a baseline, we use uncompressed streaming requiring approximately 300 megabits per second and our solution achieves similar visual quality with tiled adaptive streaming at 14 megabits per second. We also demonstrate statistically significant gains in the quality of interaction and improvements to system performance and CPU consumption with tiled adaptive streaming as compared to the more traditional network adaptive streaming.
Dick C. A. Bulterman合作论文数Centrum Wiskunde & Informatica;Department of Computer Science, Faculteit der Exacte Wetenschappen, Vrije Universiteit Amsterdam50