The goal of Mixed Reality (MR) is to achieve a seamless and realistic blending between real and virtual worlds. This requires the estimation of reflectance properties and lighting characteristics of the real scene. One of the main challenges within this task consists in recovering such properties using a single RGB-D camera. In this article, we introduce a novel framework to recover both the position and color of multiple light sources as well as the specular reflectance of real scene surfaces. This is achieved by detecting and incorporating information from both specular reflections and cast shadows. Our approach is capable of handling any textured surface and considers both static and dynamic light sources. Its effectiveness is demonstrated through a range of applications including visually-consistent mixed reality scenarios (e.g., correct real specularity removal, coherent shadows in terms of shape and intensity) and retexturing where the texture of the scene is altered whereas the incident lighting is preserved.
This article has been removed by arXiv administrators because the submitter did not have the rights to agree to the license at the time of submission
Intrinsic image decomposition describes an image based on its reflectance and shading components. In this paper we tackle the problem of estimating the diffuse reflectance from a sequence of images captured from a fixed viewpoint under various illuminations. To this end we propose a deep learning approach to avoid heuristics and strong assumptions on the reflectance prior. We compare two network architectures: one classic ‘U’ shaped Convolutional Neural Network (CNN) and a Recurrent Neural Network (RNN) composed of Convolutional Gated Recurrent Units (CGRU). We train our networks on a new dataset specifically designed for the task of intrinsic decomposition from sequences. We test our networks on MIT and BigTime datasets and outperform state-of-the-art algorithms both qualitatively and quantitatively.
In this paper, we consider the problem of estimating the 3D position and intensity of multiple light sources without using any light probe or user interaction. The proposed approach is twofold and relies on RGB-D data acquired with a low cost 3D sensor. First, we separate albedo/texture and illumination using lightness ratios between pairs of points with the same reflectance property but subject to different lighting conditions. Our selection algorithm is robust in presence of challenging textured surfaces. Then, estimated illumination ratios are integrated, at each frame, within an iterative process to recover position and intensity of light sources responsible of cast shadows. Estimated lighting characteristics are finally used to achieve realistic Augmented Reality (AR).
In this work, we consider the challenge of achieving a coherent blending between real and virtual worlds in the context of a Mixed Reality (MR) scenario. Specifically, we have designed and implemented an interactive demonstrator that shows a realistic MR application without using any light probe. The proposed system takes as input the RGB stream of the real scene, and uses these data to recover both the position and intensity of light sources. The lighting can be static and/or dynamic and the geometry of the scene can be partially altered. Our system is robust in presence of specular effects and handles both uniform and/or textured surfaces.
Illumination estimation is often used in mixed reality to re-render a scene from another point of view, to change the color/texture of an object, or to insert a virtual object consistently lit into a real video or photograph. Specifically, the estimation of a point light source is required for the shadows cast by the inserted object to be consistent with the real scene. We tackle the problem of illumination retrieval given an RGBD image of the scene as an inverse problem: we aim to find the illumination that minimizes the photometric error between the rendered image and the observation. In particular we propose a novel differentiable renderer based on the Blinn-Phong model with cast shadows. We compare our differentiable renderer to state-of-the-art methods and demonstrate its robustness to an incorrect reflectance estimation.
Photometric registration consists in blending real and virtual scenes in a visually coherent way. To achieve this goal, both reflectance and illumination properties must be estimated. These estimates are then used, within a rendering pipeline, to virtually simulate the real lighting's interaction with the scene. In this paper, we are interested in indoor scenes where light bounces off of objects with different reflective properties (diffuse and/or specular). In these scenarios, existing solutions often assume distant lighting or limit the analysis to a single specular object. We address scenes with various objects captured by a moving RGB-D camera and estimate the 3D position of light sources. Furthermore, using spatio-temporal data, our algorithm recovers dense diffuse and specular reflectance maps. Finally, using our estimates, we demonstrate photo-realistic augmentations of real scenes (virtual shadows, specular occlusions) as well as virtual spec-ular reflections on real world surfaces.
Augmented Reality (AR) scenarios aim to provide realistic blending between real world and virtual objects. A key factor for realistic AR is thus a correct illumination simulation. This consists in estimating the characteristics of real light sources and use them to model virtual lighting. In this paper, we briefly introduce a novel method for recovering both 3D position and intensity of multiple light sources using detected cast shadows. Our algorithm has been successfully tested on a set of real scenes where virtual objects have visually coherent shadows.
Technicolor has been investigating how Mixed Reality technology could impact the future of home entertainment. We have designed and implemented a system to extend a standard TV experience with AR content, using a consumer tablet or a headset. A virtual TV mosaic is displayed around the TV screen and used as a GUI to control both TV and MR content. Using this interface, the user is able to switch TV content, display meta-data in AR (subtitles, text information or program guide), enhance TV content with interactive 3D objects blended in the environment, or play a game in interaction with the real world. The interactions between the real and the virtual worlds are handled thanks to a scene analysis pre-processing stage, which provides information about both the geometry and the lighting of the real environment. The real-virtual interactions strongly contribute to reinforcement of the immersion feeling. User feedback shows that the concept is very promising.
The Extended TV application allows to enhance an audiovisual content displayed on a TV, using Mixed Reality technology. During preprocessing, the close environment of the TV is scanned using a consumer depth camera. The captured RGB-D data are analyzed, providing models for both the 3D geometry and the lighting of the real scene. During runtime, the TV is watched through a tablet, and virtual objects can apparently come out of the screen and start populating the user's environment. Virtual objects can be occluded by real objects, and virtual shadows are consistent with the real ones.
•Long-term dense motion estimation based on multi-step optical flows is addressed.•Optical flows are combined through multi-step integration and statistical selection.•An analysis of available single-reference complexity reduction schemes is provided.•Our new multi-reference frames processing reaches longer accurate displacement fields.
The acquisition of surface material properties and lighting conditions is a fundamental step for photo-realistic Augmented Reality (AR). In this paper, we present a new method for the estimation of diffuse and specular reflectance properties of indoor real static scenes. Using an RGB-D sensor, we further estimate the 3D position of light sources responsible for specular phenomena and propose a novel photometry-based classification for all the 3D points. Our algorithm allows convincing AR results such as realistic virtual shadows as well as proper illumination and specularity occlusions.
We present statistical multi-step flow, a new approach for dense motion estimation in long video sequences. Towards this goal, we propose a two-step framework including an initial dense motion candidates generation and a new iterative motion refinement stage. The first step performs a combinatorial integration of elementary optical flows combined with a statistical candidate displacement fields selection and focuses especially on reducing motion inconsistency. In the second step, the initial estimates are iteratively refined considering several motion candidates including candidates obtained from neighboring frames. For this refinement task, we introduce a new energy formulation which relies on strong temporal smoothness constraints. Experiments compare the proposed statistical multi-step flow approach to state-of-the-art methods through both quantitative assessment using the Flag benchmark dataset and qualitative assessment in the context of video editing.
We analyze the problem of how to correctly construct dense point trajectories from optical flow fields. First, we show that simple Euler integration is unavoidably inaccurate, no matter how good is the optical flow estimator. Then, an inverse integration scheme is analyzed which is more robust to bias and input noise and shows better stability properties. Our contribution is threefold: 1) a theoretical analysis that demonstrates why and in what sense inverse integration is more accurate; 2) a rich experimental validation both on synthetic and real (image) data; and 3) an algorithm for approximate online inverse integration. This new technique is precious whether one is trying to propagate information densely available on a reference frame to the other frames in the sequence or, conversely, to assign information densely over each frame by pulling it from the reference.
Accurate estimation of dense point correspondences between two distant frames of a video sequence is a challenging task. To address this problem, we present a combinatorial multistep integration procedure which allows one to obtain a large set of candidate motion fields between the two distant frames by considering multiple motion paths across the video sequence. Given this large candidate set, we propose to perform the optimal motion vector selection by combining a global optimization stage with a new statistical processing. Instead of considering a selection only based on intrinsic motion field quality and spatial regularization, the statistical processing exploits the spatial distribution of candidates and introduces an intra-candidate quality based on forward-backward consistency. Experiments evaluate the effectiveness of our method for distant motion estimation in the context of video editing.