
The 3D tele-immersive and collaborative environment provides a virtual space for the interaction of remotely dispersed users. To achieve multi-perspective rendering and realistic 3D visual effect, it is needed to transmit multiple semantically correlated 3D video streams from the source to the destination with stringent synchronization requirement. In this paper we discuss the issue of multi-stream synchronization from the general context of the multicast routing with delay and delay variation constraints, which was proved as an NP-complete problem. Then we propose a heuristic to construct a multicast network on an overlay content dissemination architecture for the solution, and show that our algorithm is asymptotically more advanced in the time complexity than existing ones. Empirical studies further verify the performance of our algorithm regarding to the temporal efficiency in various sizes of input data.
In this paper we present framework for hierarchical calibration of multi-camera based teleimmersion systems. First, we calibrate the internal camera parameters and geometry of each stereo cluster using a checkerboard. Next, we present a robust and efficient method to externally calibrate the location of the stereo clusters using virtual calibration object created by two LED markers. Our novel algorithm does not require for all the cameras to share common workspace; only pairwise overlap is required. Finally, we address geometric correspondence between several remote locations by proposing a simple calibration method.
For adults, immersive environments for collaboration and play are expensive computer-intensive novelties. For children, they are commonplace but computer-free: kindergarten classrooms and playgrounds. Kids know how to develop, and spontaneously learn from imaginary play. If children were provided with robust, intuitive, and inexpensive computer-enabled environments they would thrive, learn, and engage in creative play, not just in physical and social worlds, but within increasingly important virtual worlds. We have developed a potentially ubiquitous, robust, intuitive, and cheap computer-enabled immersive environment for children, and have designed interactive technology that can be configured to support and enhance, rather than to structure and constrain children's play.
A novel dense depth map estimation algorithm is proposed in order to meet the requirements of N-view plus N-depth representation, which is one of the standardization efforts for the upcoming 3D display technologies. Hence, extraction of multiple depth maps is achieved from multi-view video. Starting from the piecewise planarity assumption of the scene, estimation of 3D structure of the patches, obtained through color-based over-segmentation, is achieved by plane- and angle-sweeping for every view independently. Markov Random Field (MRF) modeling is utilized for each view in pixel-wise manner in order to relax and refine the estimated planar models while incorporating visibility and consistency constraints. In this algorithm, the fusion of multiple depth maps is performed by updating belief values on the observed nodes based on depth and color consistency during the refinement step. The proposed method handles untextured surfaces, as well as depth discontinuities at object boundaries, due to its initial modeling of the scene as piecewise planar regions. The experimental results illustrate reliability and the robustness of the proposed algorithm for different type of scenes.
A depth image-based rendering (DIBR) technique is one of the rendering processes of virtual views with a color image and the corresponding depth map. The most important issue of DIBR is that the virtual view has no information at newly exposed areas, so called disocclusion. The general solution is to smooth the depth map using a Gaussian smoothing filter before 3D warping. However, the filtered depth map causes geometric distortion and the depth quality is seriously degraded. Therefore, we propose a new depth map filtering algorithm to solve the disocclusion problem while maintaining the depth quality. In order to preserve the visual quality of the virtual view, we smooth the depth map with further reduced deformation. After extracting object boundaries depending on the position of the virtual view, we apply a discontinuity-adaptive smoothing filter according to the distance of the object boundary and the amount of depth discontinuities. Finally, we obtain the depth map with higher quality compared to other methods. Experimental results showed that the disocclusion is efficiently removed and the visual quality of the virtual view is maintained.
Applying recent advances in multi-imager capture and multi-projector display, we combine capabilities through the Nizza multimedia dataflow architecture to deliver low-cost wide-VGA-quality low-latency autostereoscopic 3D display of live video on a single PC. Supporting multiple users as they observe and interact against a life-sized display surface responsive to their positions, this facility will open new opportunities in mediated interaction.
Multi-camera systems have been evolving as next generation video cameras with applications including 3D reconstruction, image-based rendering, free viewpoint, and 3D TV. Quantifying visual quality of multi-camera systems is fundamental in developing these applications. In this paper, image distortion types in multi-camera systems are investigated where distortion is classified as either geometric or photometric. Examples and measurements are presented showing that single-view objective image quality measures are not adequate for perceptual assessment of multi-camera images. A new algorithm that characterizes the type of distortion in a given image captured by a multi-camera system is proposed and evaluated. The new algorithm is based on the edge intensity summation (EIS). A new EIS-based structural similarity (EISSM) quality measure is proposed. EISSM is shown to capture the perceptual fidelity that is not fully grasped by PSNR and SSIM.
Creating an immersive experience in a collaborative virtual environment, or CVE, involves more than just high definition graphics and expensive hardware. An immersive experience requires participant's input to be translated in a timely fashion to the local environment as well as to others connected across the internet. A well designed prediction algorithm will reduce the lag caused by virtual environment hardware and communications networks. The purpose of this experiment was to test the quality of an adaptive exponential smoothing algorithm at predicting human arm movement in a CVE. The results show that adaptive exponential smoothing performs as well as or better than dead reckoning at position estimation. When applied to a CVE adaptive Holt's exponential smoothing can help to reduce the overall lag of the system without being as computationally complex as many other techniques.
This paper proposes a new 3D propagation algorithm for the depth image-based rendering problem with multiple color and range cameras at arbitrary positions. The proposed algorithm efficiently renders novel images at arbitrary virtual views by propagating all available depth information from range cameras to color cameras, and then all available depth and color information from color cameras to the virtual views. Furthermore, the algorithm significantly enhances the propagated depth images by applying a new occlusion removal method and a new depth-color bilateral filtering. The paper also describes the parallelism structure of our algorithm and outlines a mapping onto massively parallel architectures such as general-purpose graphics processing units (GPGPUs). Experimental results show that the proposed algorithm provides good rendering quality while staying within computational bounds for real-time applications.
We report on a fast algorithm for the generation of cylindrical panoramic views from hand-held video sequences. Due to its high processing speed the algorithm is suited for hardware implementation into next generation video- and photo cameras. This enables the user to easily create immersive views from simple pan shots of variable quality. The individual processing steps within the algorithm are described in detail. Final results of the video to panorama conversion process along with an an outlook on how to further improve the method when implemented in consumer grade video- and photo cameras are given at the end of this paper.
We present a prototype display system for group teleconferencing that delivers the proper views to multiple local viewers in different locations. While current state-of-the-art commercial teleconferencing systems can provide highdefinition video, and proper placement and scaling of remote participants, they cannot support correct eye gaze for multiple users. A display capable of providing multiple simultaneous views to each local observer and multiple aligned cameras are required for generating a distinct and spatially-appropriate image for each local participant. If each local participant can observe the remote participants from an appropriate angle, then it becomes possible for viewers to properly identify where other participants are looking. Our system provides view-appropriate imagery for multiple users by spatially multiplexing the output of the display. We achieve this by placing a lenticular sheet over the surface of the display, which directs light from different pixels in different directions. With knowledge about the subset of the display surface each participant can see, it is possible to combine each of the remote cameras' images into a single composite image. This image, when viewed though the multiplexing layer, appears to be the appropriate camera image for each local participant. The prototype system uses a camera-based display calibration technique capable of properly evaluating which pixels are visible from an arbitrary viewpoint without a physical model of the display.
Tele-immersive systems, are growing in popularity and sophistication. They generate 3D video content in large scale, yielding challenges for executing data-mining tasks. Some of the tasks include classification of actions, recognizing and learning actor movements and so on. Fundamentally, these tasks require tagging and identifying of the features present in the tele-immersive 3D videos. We target the problem of 3D feature extraction, a relatively unexplored direction. In this paper we propose Samera, a scalable and memory-efficient feature extraction algorithm which works on short 3D video segments. The focus is on relevant portions of each frame, then uses a flow based technique across frames (in a short video segment) to extract features. Finally it is scalable, by representing the constructed feature vector as a binary vector using Bloom Filters. The results obtained from experiments performed on 3D video segments obtained from Laban Movement Analysis (LMA) show that the compression ratio achieved in Samera is 147.5 as compared to the original 3D videos.
We describe Pantheia, a system that constructs virtual models of real spaces from collections of images, through the use of visual markers that guide and constrain model construction.To create a model users simply 'mark up' the real world scene by placing preprinted markers that describe scene elements or impose semantic constraints.Users then collect still images or video of the scene.From this input, Pantheia automatically and quickly produces a model.The Pantheia system was used to produce models of two rooms that demonstrate the effectiveness of the approach.
In some visual communication applications it is not possible or even desired to aim at a photorealistic representation of the remote person. One possibility is to aim at stylized visual representations of remote persons, e.g., as avatars shown on a display device or as shadows in lighting. In this paper we introduce a system for persistent and ambient visual communication based on capture, transmission, and rendering of 3D shadow representations of users. The shape of a person is captured using a distributed camera array, compressed, and transmitted over the network. In the receiving end the shape is projected as a shadow on a surface using a lighting device. We demonstrate that the 3D representation of the shape makes it possible to control the 2D visualization at the receiving end in many interesting ways. For example, when controlled by tracking of the observing user the shadow may create a visual illusion of a 3D shape on the wall.
The European FP7 project 3DPresence is developing a multiparty, high-end 3D videoconferencing concept that tackles the problem of transmitting the feeling of physical presence in real-time to multiple remote locations in a transparent and natural way. Traditional set-top camera video-conferencing systems still fail to meet the 'telepresence challenge' of providing a viable alternative for physical business travel, which is nowadays characterized by unacceptable delays, costs, inconvenience, and an increasingly large ecological footprint. Even recent high-end commercial solutions, while partially removing some of these traditional shortcomings, still present the problems of not scaling easily, expensive implementations, not utilizing 3D life-sized representations of the remote participants and addressing only eye contact and gesture-based interactions in very limited ways. One of many challenges in this project is to calculate depth information for many different views in order to synthesise novel views to provide eye contact. In this paper, we present a multi-baseline disparity fusion scheme for improved real-time disparity map estimation. The advantages and disadvantages of different configurations are discussed and theoretical considerations are presented regarding disparity resolution and baseline. These observations together with experimental investigations lead to a multi-baseline configuration that allows taking advantage of small and wide baseline stereo camera as well as trifocal camera configurations.
Humans are spending an increasing amount of time in tele-immersive environments interacting with avatars or virtual human bodies. Additionally, human behavior and cognition are affected by experiences in tele-immersive environments. Although there is substantial psychological work surrounding the notion of morality, there is little work that examines the interplay of immersive digital environments and the moral identity of the digital medium user. We conducted a study to explore how participants' moral behaviors and self-ratings of morality changed after immersion in either a moral or immoral tele-immersive environment. Results revealed that participants who witnessed the immoral scenarios felt and acted more immoral than participants in the moral scenario condition. These findings have important implications for understanding the effects of digital media as well as for the study of the psychological construct of moral identity.
With the advent of virtual spaces, there has been a need to integrate physical world with virtual spaces.The integration can be achieved by real-time 3D imaging using stereo cameras followed by fusion of virtual and physical space information.Systems that enable such information fusions over several geographically distributed locations are called tele-immersive and should be easily deployed.The optimal placement of 3D cameras becomes the key to achieving high quality 3D information about physical spaces.In this paper, we present an optimization framework for automating the placement of multiple stereo cameras in an application specific manner.The framework eliminates ad-hoc experimentations and sub-optimal camera placements for end applications by running our simulation code.The camera placement problem is formulated as optimization problem over continuous physical space with the objective function based on 3D information error and a set of constraints that generalize application specific requirements.The novelty of our work lies in developing the theoretical optimization framework under spatially varying resolution requirements and in demonstrating improved camera placements with our framework in comparison with other placement techniques.
3D video can be represented by color 2D video sequence accompanied by gray-scale dense depth map (depth image) sequence. In this paper, we describe a novel method for intraframe compression of the depth modality of such representation. Our method takes into account specific features of depth images, i.e. the presence of large smooth regions delineated by sharp discontinuities (edges). For such images, conventional transform-based coders produce undesirable artifacts which impede the subsequent rendering of virtual views. Instead of block transforms, the proposed method employs horizontal-vertical anisotropic partition scheme which yields a tree-structured decomposition of non-overlapping rectangular blocks adapted to the depth map content. Each block in the decomposition is approximated by plane described by the block corner pixels. The codestream consists of the coded partition scheme and the coded error of prediction of quantized corner pixels. The scheme substantially reduces the amount of artifacts around edges and yields an improvement of several dB in PSNR for typical compression ratios compared with the best transform-based coders. Other advantages of the designed coder are its simplicity and fast decompression, and the possibility to control the ratedistortion performance.
3D video is an emerging technology that promises immersive experiences in a truly seamless environment. In this paper, we present a new proxy-based framework to extend the 3D video experience to mobile devices. The framework has two major features. First, it allows audience to use mobile devices to change rendering viewpoints and interact with 3D video, and it supports different interaction devices to collaborate with each other. Second, the proxy compresses the 3D video streams before broadcasting to mobile devices for display. A novel view-dependent real-time 3D video compression scheme is implemented to reduce the requirements of both transmission bandwidth and rendering computation on mobile devices. The system is implemented on different mobile platforms and our experiments indicate a promising future of 3D video on mobile devices.
The emergence of force feedback haptic devices that can remotely interact with virtual environments presents a number of challenges to the underlying networks that have to support their interactions. One important issue concerns the characterisation of haptic traffic, particularly whenever multiple users remotely interact over a network such as the Internet. Previous research has characterised the traffic produced by single contact-point haptic devices when remotely interacting with a distributed haptic virtual environments (DHVEs). The research presented in this paper extends this work to consider the more complex traffic produced by haptic devices with multiple contact points, whenever interacting remotely with virtual environments. Such devices produce a rich mixture of different traffic streams that are interdependent but also characterised by the interactions of the individual users. The aim of the work presented here is to characterise the traffic generated by multi-point DHVE network connections. The approach taken develops an analytical model of DHVE traffic based on empirical measurements. Suitable probability distributions models are subsequently derived for each type of traffic. The results show that each traffic type exhibits either a Normal or a Weibull distribution. The results permit the development of a multi-contact point haptic traffic generator model which can then be used by simulation and analytical studies in order to examine how such interactive applications can be transmitted over different network situations and topologies.