Tele-immersive systems development is always driven as well as restricted by the available immersive technology. Hence, existing such systems are described mainly from a technological point of view; their conceptual description is usually limited to the description of a scenario that is implementable with or circumvents the limitations of the chosen technology. This focus on technology makes it difficult to compare systems’ concepts; moreover, it has led to different views on tele-immersion in different fields, such as remotely controlled robots, immersive video conferencing, and tele-collaboration. In this work, we give a general, structured principle to describe the conceptual part of any tele-immersion system. This principle naturally unifies the different views on tele-immersion. Our idea is based on the insight that, in order to be general, immersion must be described separately for each direction of communication. We characterize communication between locations using a graph; for each directed edge of this graph, we describe immersion as operations on volumes. Using this principle, we define a typology, which enables the comparison and enumeration of tele-immersion concepts. We apply this typology to survey the concepts of existing tele-immersion systems and thereby demonstrate how three well-known tele-immersive scenarios—Marvin Minsky’s tele-operated robot, the Office of the Future, and the asymmetric Beaming scenario—integrate naturally. We show how the general principle can be utilized conveniently to grasp conceptual ideas in tele-immersion, such as direct interaction, locational presence, spatial consistency, symmetries, and self-inclusion.
Telepresence systems use 3D techniques to create a more natural human-centered communication over long distances. This work concentrates on the analysis of latency in telepresence systems where acquisition and rendering are distributed. Keeping latency low is important to immerse users in the virtual environment. To better understand latency problems and to identify the source of such latency, we focus on the decomposition of system latency into sub-latencies. We contribute a model of latency and show how it can be used to estimate latencies in a complex telepresence dataflow network. To compare the estimates with real latencies in our prototype, we modify two common latency measurement methods. This presented methodology enables the developer to optimize the design, find implementation issues and gain deeper knowledge about specific sources of latency.
Depth cameras are increasingly used for tasks such as 3-D reconstruction, user pose estimation, and humancomputer interaction. Depth-camera systems comprising multiple depth sensors require careful calibration. In addition to conventional 2-D camera calibration, depth correction for each individual device is necessary. In this paper, we present a new way of solving the multi depth-camera calibration problem. Our main contribution is a novel depth correction approach which supports the generation of a 3-D lookup table by incorporating an optical marker-based tracking system. We verify our approach for the Microsoft Kinect and for the MESA SwissRanger4000 time-of-flight camera.
Life-size high-resolution telepresence systems, used for remote collaboration, face the problem of transmitting huge data from multiple viewpoints. We present different strategies focusing on efficient camera selection and acquisition method to discard part of image data for transmission as a preprocess to classical video compression schemes. At the the receiver site, part of the omitted data can be restored by means of super-resolution methods.
We present a distributed dataflow-based camera library design. It is tailored to the special needs of distributed acquisition and rendering telepresence systems with inhomogeneous camera arrays. We discuss a variety of design decisions: choice of dataflow concept, memory management, camera interface abstraction and important built-in functionality.
Systems for multi-camera telepresence utilizing Large High-Resolution Displays (LHRDs) face the problem of generating large amounts of dynamic data. These data have to be processed at the local capture site, aggregated from different nodes locally, possibly processed again, transmitted to a remote site, and processed further. Since resources are limited a solution to reduce the amount of data is essential to build such systems. In the early stages of an image processing pipeline traditional video compression is not an option as following image processing stages would have to decompress these images again. In this paper we present our approaches to bandwidth optimization in such setups with special focus on the local transmission occurring in distributed acquisition setups. We show different strategies and give an overview of where these are applicable. The results we achieved implementing Dynamic Frustum Selection in our prototype setup are presented and analyzed. We provide the reader with a guide for design decisions when implementing such or comparable systems.
Systems for immersive videoconferencing utilizing LHRDs face the problem of generating large amounts of data. A solution to reduce this data is essential to realize such systems. We present different possible approaches to do this and give an overview where these are applicable. This provides the reader with a first guide for design decisions when implementing such systems.
Segmenting foreground from background automatically is an active field of research. The graph cut approach is one of the promising methods to solve this problem. This approach requires that the weights of the graph are chosen optimally in order to obtain a good segmentation. We address this challenge focusing on the automatic segmentation of wood log images. We present a novel method based on density estimation to obtain information about both foreground and background. With this information the weights in the graph cut method can be set automatically. In order to validate our results, we use four different methods to set these weights. We show that of these approaches, our new method obtains the best results.
We present a novel tele-presence approach that extends the window metaphor by combining large high-resolution LCD walls with multi-camera 3D video. We propose to integrate an array of cameras into the bezels of the wall to support flexible camera placement for optimized video acquisition. The users's 3D video representation combined with the high-resolution LCD wall provides local and remote users with a shared virtual space in an extended life-size window metaphor. We discuss important system design aspects such as camera placement strategies, resolution, field of view, and dynamic camera selection for different 3D video reconstruction approaches, such as stereo and visual hulls. Finally, we describe our current prototype system based on the design guidelines described in this paper.
The automatic extraction of foreground objects from the background is a well known problem. Much research has been done to solve the foreground/background segmentation with graph cuts. The major challenge is to determine the weights of the graph in order to obtain a good segmentation. In this paper we address this problem with a focus on the automatic segmentation of wood logs. We introduce a new solution to get information about foreground and background. This information is used to set the weights of the graph cut method. We compare four different methods to set these weights and show that the best results are obtained with our novel method, which is based on density estimation.