We present Mover360, a controllable object manipulation framework for 360° images. Unlike perspective images, 360° images in equirectangular projection (ERP) exhibit horizontal wrap-around, latitude-dependent distortion, and global scene continuity, which makes object-level edits difficult for existing perspective editors to produce and for users to specify. To address this, Mover360 centers on object Translation (relocating a specified object within an existing panorama) while supporting reference-guided Insert and Remove as auxiliary tasks. Its interface unifies point-, bbox-, and mask-guided control by encoding each task into a fixed prompt and a compact, ERP-aligned instruction map. In the default point mode, a single click relocates an object, allowing the model to infer a plausible size, support, and illumination using panoramic context and an auxiliary depth condition. Structurally, Mover360 is a lightweight adaptation of a pretrained diffusion transformer. To generate paired supervision, we construct a UE5 data-generation pipeline with surface-aware object placement and randomized illumination, yielding large-scale paired data and a dual-domain benchmark of synthetic and real panoramas with ground truth for all three tasks. Across both test domains and two evaluation protocols, Mover360 outperforms strong baselines for perspective editing, insertion, and inpainting in reconstruction fidelity, semantic consistency, and distributional quality. Code and our benchmark dataset are available at https://zhonghaoyi.github.io/Mover360/.
While instruction-based image editing is emerging, extending it to 360° panorama introduces additional challenges. Existing methods often produce implausible results in both equirectangular projections (ERP) and perspective views. To address these limitations, we propose SE360, a novel framework for multi-condition guided object editing in 360° panoramas. At its core is a novel coarse-to-fine autonomous data generation pipeline without manual intervention. This pipeline leverages a Vision-Language Model (VLM) and adaptive projection adjustment for hierarchical analysis, ensuring the holistic segmentation of objects and their physical context. The resulting data pairs are both semantically meaningful and geometrically consistent, even when sourced from unlabeled panoramas. Furthermore, we introduce a cost-effective, two-stage data refinement strategy to improve data realism and mitigate model overfitting to erasing artifacts. Based on the constructed dataset, we train a Transformer-based diffusion model to allow flexible object editing guided by text, mask, or reference image in 360° panoramas. Our experiments demonstrate that our method outperforms existing methods in both visual quality and semantic accuracy.
360° images, when paired with virtual reality (VR) headsets, offer an immersive viewing experience of real-world environments. The integration of mixed reality (MR) within 360° videos further enhances users’ telepresence by blending virtual assets into the real environment. This paper investigates how high-fidelity visual appearance and physical interaction between virtual and real objects affect user perception in MR 360° videos. We focus on two key features, occlusion and collision, that enable high fidelity interaction with virtual objects within a complex mixed reality 360° scene. To study their impact on visual realism and users’ sense of presence, we conduct a user study using simulated 360° videos with pixel-aligned depth data.
In this article, we propose an AR system that facilitates a user's natural interaction with virtual objects in an augmented reality environment. The system consists of three modules: human pose and shape estimation, camera-space calibration, and physics simulation. The first module estimates a user's 3D pose and shape from a single RGB video stream, thereby reducing the system setup cost and broadening potential applications. The camera-space calibration module estimates the user's camera-space position to align the user with the input RGB image. The physics simulation enables seamless and physically natural interaction with virtual objects. Two prototyping applications built upon the system prove an enhancement in the quality of interaction, fostering a more immersive and intuitive user experience.
This paper presents a pipeline that estimates a user’s 3D pose and shape to facilitate a user’s full-body interaction with virtual objects in a mixed reality environment. The usability and effectiveness of the pipeline are demonstrated through a user study.
Neural Radiance Fields (NeRF) have demonstrated promising results in synthesizing novel view images from a set of unconstrained captured scenes. One important extension of NeRF is using it on non-rigid reconstruction. Although previous NeRF-based methods for dynamic scene reconstruction have presented visually appealing results, they still often show visual artifacts such as blurry or incorrect geometry of an object. One of the causes is that previous work performs reconstruction directly on the entire video sequence. The global temporal information over the video sequence introduces noise to the network, often leading to a non-optimal canonical space representation of the dynamic scene. In this paper, we present Local Temporal (LT) NeRF, a method to synthesize novel views of dynamic scenes using local temporal priors. Our novel LT module provides the local temporal priors using multi-view stereo sampling, and improves the deformation field reconstruction and hyper-space encoding. Our novel loss functions further supervise the NeRF for better optimization. We evaluate our method with dynamic scenes captured from monocular videos, outperforming the state-of-the-art.
Removing undesired shading from human images is crucial in supporting various real-world applications. While recent advancements in deep learning-based methods show promise in addressing this challenge, there persists a struggle to accurately separate texture from shading, which often results in unresolved shading artifacts and altered texture patterns. This issue is exacerbated by dataset limitations, such as the lack of diverse real-world clothing styles in realistic datasets and oversimplified assumptions about human reflectance and illumination environments. To address this problem, our paper introduces a novel semi-supervised deep learning method to effectively assemble both real and synthetic data for better disentanglement of texture and shading. We present a global sparsity constraint designed on both labeled and unlabeled data to minimize color variations in the inferred shading map, enhancing texture recovery. By applying this constraint, our method demonstrates improved handling of a broad range of fashion-related textures in the real-world test. Additionally, we address the disparity between real and synthetic data with a novel domain adaptation module to realize effective transfer from synthetic to real images. This module is designed based on the insights of gamma correction, and demonstrates improved shadow removal in real-world images. By integrating these methods, our approach achieves state-of-the-art results, reducing unwanted shading artifacts while maintaining the integrity of underlying textures in real-world scenarios.
360 degrees images offer panoramic views of captured environments, placing users within an egocentric perspective. While users can freely rotate their viewpoint, they don't experience 6-DoF navigation with translational movement. In this research, we introduce Avatar360, a novel method to elicit 6-DoF perception in 360 degrees panoramas, using avatar-assisted navigation combined with an exocentric view of the 360 degrees panorama. We seamlessly integrate a 3D avatar into 360 degrees panoramas, allowing users to navigate a 3D virtual landscape congruent with the 360 degrees background. By aligning the exocentric perspective of the 360 degrees panorama with the avatar's movements, we replicate a sensation of 6-DoF navigation in 360 degrees panoramas. We explore mechanisms for simultaneous avatar and viewpoint controls, as well as procedures for transitions between spatially connected 360 degrees panoramas. A user study was conducted to assess the perception of 6-DoF navigation in 360 degrees panoramas via a 3D avatar, evaluating users' sense of movement, disorientation, and presence. We also gained insight into perspective view controls and transition techniques between panoramas. Statistical analysis shows avatar-assisted navigation elicits a user's sense of movement within 360 degrees panoramas. Our results also provide guidelines for effective view control and transition strategies in avatar-assisted 360 degrees navigation.
Augmented telepresence provides rich communication for people at a distance with interactive blended information between the virtual and real world [Rhee et al. 2017, 2020; Young et al. 2022]. We push the boundaries of augmented telepresence with a novel live media technology, including live capturing, modeling, blending, and interactive effects (IFX) to augment telepresence. Using our technology, people at a distance can connect and communicate with creative storytelling, augmented with novel IFX. We achieve this with the following breakthroughs: 1) digitizing remote spaces and people in real-time, 2) transmitting digitized information across a network, 3) augmenting remote telepresence using real-time visual effects and interactive storytelling with live-blending of 3D virtual assets into the digitized real-world. In this presentation, we will unveil several new technologies and novel IFX that can enrich telepresence, including: • Real-time 360° RGBD video capturing: we will demonstrate capturing 360° RGBD videos using a 360° RGB camera and LiDAR sensor, including synchronization between the RGB and depth streams as well as depth map generation. • IFX with live RGBD videos: we will demonstrate real-time blending of 3D virtual objects into the live 360° RGBD videos, showcasing real-time occlusion and collision handling. • 6-degrees of freedom (DoF) tele-movement: we introduce our recent research [Chen et al. 2022] for volumetric environment capturing and 6-DoF navigation. We will demonstrate real-time navigation (movement and rotation) in captured real surroundings (beyond room scales). We will showcase applications (Figure 1) where we can virtually teleport to and explore within a live stream of the augmented real world and communicate remotely with live IFX.
Six degrees-of-freedom (6-DoF) video provides telepresence by enabling users to move around in the captured scene with a wide field of regard. Compared to methods requiring sophisticated camera setups, the image-based rendering method based on photogrammetry can work with images captured with any poses, which is more suitable for casual users. However, existing image-based rendering methods are based on perspective images. When used to reconstruct 6-DoF views, it often requires capturing hundreds of images, making data capture a tedious and time-consuming process. In contrast to traditional perspective images, 360° images capture the entire surrounding view in a single shot, thus, providing a faster capturing process for 6-DoF view reconstruction. This article presents a novel method to provide 6-DoF experiences over a wide area using an unstructured collection of 360° panoramas captured by a conventional 360° camera. Our method consists of 360° data capturing, novel depth estimation to produce a high-quality spherical depth panorama, and high-fidelity free-viewpoint generation. We compared our method against state-of-the-art methods, using data captured in various environments. Our method shows better visual quality and robustness in the tested scenes.
We developed the Motion-Simulation Platform, a platform running within a game engine that is able to extract both RGB imagery and the corresponding intrinsic motion data (i.e., motion field). This is useful for motion-related computer vision tasks where large amounts of intrinsic motion data are required to train a model. We describe the implementation and design details of the Motion-Simulation Platform. The platform is extendable, such that any scene developed within the game engine is able to take advantage of the motion data extraction tools. We also provide both user and AI-bot controlled navigation, enabling user-driven input and mass automation of motion data collection.
We present MRMAC, a Mixed Reality Multi-user Asymmetric Collaboration system that allows remote users to teleport virtually into a real-world collaboration space to communicate and collaborate with local users. Our system enables telepresence for remote users by live-streaming the physical environment of local users using a 360° camera while blending 3D virtual assets into the mixed-reality collaboration space. Our novel client-server architecture enables asymmetric collaboration for multiple AR and VR users and incorporates avatars, view controls, as well as synchronized low-latency audio, video, and asset streaming. We evaluated our implementation with two baseline conditions: conventional 2D and standard 360° videoconferencing. Results show that MRMAC outperformed both baselines in inducing a sense of presence, improving task performance, usability, and overall user preference, demonstrating its potential for immersive multi-user telecollaboration.
This study aims to use deep learning to recover the diffuse albedo of human images captured under a wide range of real-world lighting conditions. A key challenge here is the wide variety of textures found in full-body human images. While some aspects like skin color have a limited color range, clothing and accessories display a broad spectrum of colors and textures. As a result, creating a comprehensive dataset with accurate labels is unfeasible. To address this, we propose a data augmentation method that involves applying color-shifts to various semantic regions within our training images, all while maintaining realistic appearance. This process is accomplished by initially segmenting the ground-truth albedos into their respective components (e.g., pants, shirt, hair, etc.) using a pre-trained human parsing network. Then, we adjust their hue and intensity channels using randomly chosen values from a carefully defined distribution. Our results show significant improvements in albedo recovery, especially in clothing areas, and better performance with underrepresented skin tones.
We present a novel live platform enhancing stage performances with real-time visual effects. Our demo showcases real-time 3D modeling, rendering and blending of assets, and interaction between real and virtual performers. We demonstrate our platform’s capabilities with a mixed reality performance featuring virtual and real actors engaged with in-person audiences.
This paper presents a novel solution for estimating simulator sickness in HMDs using machine learning and 3D motion data, informed by user-labeled simulator sickness data and user analysis. We conducted a novel VR user study, which decomposed motion data and used an instant dial-based sickness scoring mechanism. We were able to emulate typical VR usage and collect user simulator sickness scores. Our user analysis shows that translation and rotation differently impact user simulator sickness in HMDs. In addition, users' demographic information and self-assessed simulator sickness susceptibility data are collected and show some indication of potential simulator sickness. Guided by the findings from the user study, we developed a novel deep learning-based solution to better estimate simulator sickness with decomposed 3D motion features and user profile information. The model was trained and tested using the 3D motion dataset with user-labeled simulator sickness and profiles collected from the user study. The results show higher estimation accuracy when using the 3D motion data compared with methods based on optical flow extracted from the recorded video, as well as improved accuracy when decomposing the motion data and incorporating user profile information.
Six degrees-of-freedom (6-DoF) video provides telepresence by enabling users to move around in the captured scene with a wide field of regard. Compared to methods requiring sophisticated camera setups, the image-based rendering method based on photogrammetry can work with images captured with any poses, which is more suitable for casual users. However, existing image-based rendering methods are based on perspective images. When used to reconstruct 6-DoF views, it often requires capturing hundreds of images, making data capture a tedious and time-consuming process. In contrast to traditional perspective images, 360° images capture the entire surrounding view in a single shot, thus, providing a faster capturing process for 6-DoF view reconstruction. This article presents a novel method to provide 6-DoF experiences over a wide area using an unstructured collection of 360° panoramas captured by a conventional 360° camera. Our method consists of 360° data capturing, novel depth estimation to produce a high-quality spherical depth panorama, and high-fidelity free-viewpoint generation. We compared our method against state-of-the-art methods, using data captured in various environments. Our method shows better visual quality and robustness in the tested scenes.
Simulating marine sponge growth helps marine biologists analyze, measure, and predict the effects that the marine environment has on marine sponges, and vice versa. This paper describes a way to simulate and grow geometric models of the marine sponge Crella incrustans while considering environmental factors including fluid flow and nutrients. The simulation improves upon prior work by changing the skeletal architecture of the sponge in the growth model to better suit the structure of Crella incrustans. The change in skeletal architecture and other simulation parameters are then evaluated qualitatively against photos of a real-life Crella incrustans sponge. The results support the hypothesis that changing the skeletal architecture from radiate accretive to Halichondrid produces a sponge model which is closer in resemblance to Crella incrustans than the prior work.
Radiance maps (RM) are used for capturing the lighting properties of real-world environments. Databases of RMs are useful for various rendering applications such as look development, live action composition, mixed reality, and machine learning. Such databases are not useful if they cannot be organized in a meaningful way. To address this, we introduce the illumination space, a feature space that arranges RM databases based on illumination properties. Our method is motivated by how the RM illuminates the scene as opposed to describing the textural content of the RM. We avoid manual labeling by automatically extracting features from an RM that provides a concise and semantically meaningful representation of its typical lighting effects. We also introduce ‘Illumination Browser’, a user interface (UI) that visualizes the illumination space alongside a real-time preview renderer to enable intuitive browsing for artists. This is made possible with the following contributions: a method to automatically extract a small set of dominant and ambient lighting properties from RMs, a low-dimensional (5D) light feature vector summarizing these properties to form the illumination space, and a UI that effectively utilizes the illumination space.
We present a deep neural network for removing undesirable shading features from an unconstrained portrait image, recovering the underlying texture. Our training scheme incorporates three regularization strategies: masked loss, to emphasize high-frequency shading features; soft-shadow loss, which improves sensitivity to subtle changes in lighting; and shading-offset estimation, to supervise separation of shading and texture. Our method demonstrates improved delighting quality and generalization when compared with the state-of-the-art. We further demonstrate how our delighting method can enhance the performance of light-sensitive computer vision tasks such as face relighting and semantic parsing, allowing them to handle extreme lighting conditions.
I will talk about the importance of lighting in mixed reality rendering. From perceptually-based lighting to machine learning-driven light estimation. I also discuss using these research outcomes to inform and create lighting tools for VFX artists. Finally, I generalize the idea of virtual lighting to future research directions involving the creation of digital twins of physical objects and environments.