In low-light scenarios such as nighttime inspection and post-disaster search and rescue, robotic-arm teleoperation is often hindered by insufficient illumination, which leads to missing geometric details and degraded textures in the scene. This severely limits the accuracy and reliability of 3D reconstruction and, in turn, compromises the execution of teleoperation tasks. Existing low-light enhancement methods mostly focus on improving the perceptual quality of single-view images, while lacking explicit modeling of intrinsic multi-view consistency. As a result, downstream 3D reconstruction suffers from bottlenecks such as failed feature matching and cross-view artifacts. To address these challenges, this paper proposes a unified reconstruction framework that integrates multi-view consistent diffusion enhancement with 3D Gaussian Splatting (3DGS). Specifically, a joint optimization mechanism is introduced into the diffusion model by incorporating 2D reconstruction regularization and a cross-view consistency loss. This design enhances image illumination and recovers scene details while strictly enforcing high consistency in photometric features across views. Furthermore, the enhanced multi-view image sequences are leveraged to achieve efficient and high-fidelity scene-level 3D Gaussian geometric representation. Finally, the reconstructed 3D scene representation is seamlessly integrated into a virtual reality platform to support immersive interaction and teleoperation in complex environments. Experimental results demonstrate that, in real-world low-light scenes, the proposed method not only significantly improves image visual quality and 3D reconstruction accuracy, but also provides precise visual perception for robotic-arm teleoperation and grasping tasks, thereby validating its strong application potential in low-light environments.
Redirected walking (RDW) subtly adjusts the user's visual perspective on head-mounted displays during natural walking to reduce forced resets, thus enlarging the size of the virtual environment that can be explored beyond that of the physical environment. Alignment-based RDW controllers aim to minimize spatial discrepancies by optimizing the alignment between the user's physical and virtual environments. We introduce a novel alignment-based method that dynamically calculates mapping functions between physical and virtual geometries to enhance the algorithm's awareness of the RDW environments. To achieve this, we first construct an abstract model defining a mapping function between physical and virtual geometries and establish feasibility constraints in differential form. We then concretize this mapping, optimize it, and develop a practical implementation for dynamic geometric mapping in RDW. Our approach distinguishes itself by determining dense spatial mappings around the user, rather than aligning environments according to limited metrics. Through extensive testing, our algorithm has proven to markedly decrease reset incidents in natural walking, surpassing existing RDW controllers. The introduction of dynamic geometric mapping provides a fresh perspective, contributing significant insights and advancing the field.
The growing demand for diverse and realistic character animations in video games and films has driven the development of natural language-controlled motion generation systems. While recent advances in text-driven 3D human motion synthesis have made significant progress, generating realistic multi-person interactions remains a major challenge. Existing methods, such as denoising diffusion models and autoregressive frameworks, have explored interaction dynamics using attention mechanisms and causal modeling. However, they consistently overlook a critical physical constraint: the explicit spatial distance between interacting body parts, which is essential for producing semantically accurate and physically plausible interactions. To address this limitation, we propose InterDist, a novel masked generative Transformer model operating in a discrete state space. Our key idea is to decompose two-person motion into three components: two independent, interaction-agnostic single-person motion sequences and a separate interaction distance sequence. This formulation enables direct learning of both individual motion and dynamic spatial relationships from text prompts. We implement this via a VQ-VAE that jointly encodes independent motions and relative distances into discrete codebooks, followed by a bidirectional masked generative Transformer that models their joint distribution conditioned on text. To better align motion and language, we also introduce a cross-modal interaction module to enhance text-motion association. Our approach ensures the generated motions exhibit both semantic alignment with textual descriptions and preserving plausible inter-character distances, setting a new benchmark for text-driven multi-person interaction generation.
Recent advances in image-to-video generation have improved visual realism, making physically grounded and controllable dynamics an important step toward future world simulation. Current models often generate plausible motion, but it is not reliably governed by explicit physical causes, and instance-level constraints can leak or become entangled in multi-object interactions. We attribute this gap to two missing pieces: large-scale, fine-grained physical parameterization, and model designs that correctly bind physical attributes to instances and emphasize dynamics over appearance. To bridge this gap, we introduce PhyParam-Dataset, an interaction-centric collection of 130K physically simulated videos with dense physical parameterization, including force vectors, object material properties, and environmental constants across five representative rigid-body motion types. Built on this data, we present PhyParam, a physics-guided image-to-video diffusion model that conditions on object-level forces, masses, friction, restitution, and scene-level gravity via a lightweight physical-attention routing mechanism, and further improves motion learning with semantic-structural feature-space supervision. We also establish PhyParam-Bench, a benchmark for physical-law consistency in image-to-video generation, with a multi-level protocol evaluating temporal dynamics, spatial stability, and semantic–physical alignment. Experiments show that PhyParam improves physical consistency while maintaining high visual fidelity, advancing explicit rigid-body physical-parameter control for image-to-video generation. We will publicly release the dataset, benchmark, and code to support future research.
Reconstructing a 4D spatio-temporal representation of a dynamic scene from monocular video is a fundamental yet highly challenging problem in computer vision and computer graphics. Recent advances in 3D Gaussian Splatting (3DGS) for static scenes have significantly improved rendering efficiency and visual fidelity. However, extending 3DGS to dynamic scenes from single-view input remains difficult, as the lack of dynamic point cloud supervision often hinders the accurate modelling of moving objects, leading to suboptimal performance. In this paper, we introduce SD-4DGS, a novel 4D Gaussian Splatting (4DGS) method with spatial densification, specifically designed for dynamic scene reconstruction from monocular video. SD-4DGS features a fast spatial densification strategy that converts sparse point clouds into dense representations to better capture high-frequency geometry and textures. In addition, a sliding window motion regularizer, together with a stage-wise training schedule, progressively refines appearance and motion. Experiments show that our method significantly outperforms existing methods in terms of modelling quality across various datasets.
With the development of deep neural networks and differentiable rendering techniques, neural rendering methods, represented by Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS), have made significant progress. NeRF represents a 3D scene by encoding the appearance and geometry of the scene through neural networks, which are conditioned on both position and viewpoint. In contrast, 3DGS models the scene with a set of Gaussian ellipsoids, allowing for efficient rendering through the rasterization of these ellipsoids into images. However, both two methods are limited to representing static scenes. The rendering and reconstruction of dynamic scenes are critical in virtual reality and computer graphics. As such, extending neural rendering methods from static to dynamic scenes has become an important area of research. This survey organizes dynamic scene rendering methods based on NeRF and 3DGS and categorizes them according to different motion representations. Furthermore, it highlights the relevant applications of dynamic scene rendering, such as autonomous driving, digital humans, and 4D generation. Finally, we summarize the development of dynamic scene rendering and discuss the remaining limitations and open challenges.
Personal space plays a critical role in user privacy, safety, and comfort within social virtual reality (VR). While prior work has examined its spatial properties and protective mechanisms, existing systems typically rely on simple geometric boundaries or basic visual indicators that lack expressive and intuitive representation of personal space. We present MySpace, a toolkit that enables users to customize their personal space using expressive, metaphor-based visualizations in immersive environments. MySpace introduces a design framework that combines metaphor and visibility modulation. Users can choose from three metaphor categories—virtual forms (area, bubble), everyday objects (fence, umbrella), and natural phenomena (fog, falling leaves)—each offering a distinct representation style and emotional resonance. Visibility is modulated through properties such as opacity, motion, and spatial extent to suit individual preferences. Implemented in Unity with real-time customization capabilities, our demo showcases multi-user scenarios where users can dynamically adjust their personal space representations. By allowing users to personalize how their personal space is visually expressed and perceived, MySpace supports intuitive communication of spatial boundaries, promotes social awareness, and enables safer, more comfortable interactions in shared virtual environments.
Social virtual reality enables multi-user co-presence and collaboration but introduces privacy challenges such as personal space intrusion and unwanted interruptions. Teleportation negotiation techniques help address these issues by allowing users to define teleportation-permitted zone, maintaining spatial boundaries and comfort. However, existing methods primarily focus on forward view and often require physical rotation to monitor and respond to requests originating from behind. This can disrupt immersion and reduce social presence.To better understand these challenges, we first conducted a preliminary study to identify users’ needs for rear-space awareness during teleportation negotiation. Based on the findings, we designed two rear-awareness negotiation techniques, Window negotiation and MiniMap negotiation. These techniques display rear-space information within the forward view and allow direct interaction without excessive head movement. In a within-subjects study with 16 participants in a virtual museum, we compared these methods against a baseline front-facing approach. Results showed that MiniMap was the preferred technique, significantly improving spatial awareness, usability, and user comfort. Our findings emphasize the importance of integrating rear-space awareness in social VR negotiation systems to enhance interaction efficiency, comfort, and immersion.
Video amodal completion (VAC) aims to mimic the human brain's ability to implicitly perceive the complete appearance of partially occluded objects, thereby facilitating recognition and understanding. Existing VAC methods finetune video generation models on custom datasets, yet these datasets often have unrealistic distributions and small scales due to the challenges of collecting real amodal data and thus limit their performance and generalization.To address this, we utilize pre-trained image inpainting models for VAC and introduce in-context (IC) learning to enhance inter-frame consistency. However, despite the satisfactory performance of DiT-based IC Learning in generation tasks, task-agnostic global information often utilizes irrelevant scene information, resulting in completion failures when applied to amodal completion task. Additionally, IC Learning faces a cold-start problem with the exemplar construction. To this end, we propose a consistency video amodal completion with rectified in-context exemplar guidance. Specifically, we introduce rectified exemplar-guided completion by adjusting the attention weights of exemplar image relative to the target images for consistent completion, and adopt a dual-frame calibrated exemplar rectification to tackle the cold-start issue.Quantitative and qualitative experiments demonstrate that our method outperforms SOTAs, especially in terms of generalization and robustness on uncommon data and under severe occlusion.
Geographically dispersed users often rely on virtual avatars as intermediaries to facilitate interactive communication and collaboration. However, existing methods for augmented reality (AR) telepresence applications exhibit limitations, including restricted movement within confined sub-areas, lack of smooth transitions, and the necessity for manually establishing object mapping between dissimilar environments. We present a novel interactive AR framework for virtual avatar locomotion adaption while preserving semantic coherence across dissimilar indoor scenes. Initially, we conduct a preliminary user study to identify key attributes influencing preferred avatar movement. These attributes are quantified as features, and a dataset of user annotations on avatar movements is created. Based on the user interaction and scene configurations, we employ a deep reinforcement learning neural network to guide the avatar to the ideal position while maximizing semantic coherence. We validate our proposed framework through simulations and user studies by implementing an AR-based 3D telepresence prototype, demonstrating the efficacy of our framework in conveying user intentions across dissimilar environments, enabling natural and immersive 3D telepresence interactions.
As hardware and information technology continually advance, virtual reality (VR) has permeated numerous sectors, with applications becoming increasingly sophisticated. The evolution of VR systems has expanded from the seminal 3I characteristics—immersion, interaction, and imagination—to encompass 6I, incorporating intelligentization, interconnection, and iteration. The intelligentization of VR technology, an inevitable progression, has garnered heightened interest, particularly fueled by the emergence of artificial intelligence (AI) models and techniques like neural radiance fields, 3D Gaussian Splatting, neural rendering, generative adversarial networks, diffusion models, and large language models, which significantly propel the development of VR’s core and pivotal technologies. This survey offers a comprehensive assessment of these pivotal VR technologies, harnessing the latest AI advancements, aiming to provide fresh perspectives and assist new researchers in staying abreast of groundbreaking work. We commence by detailing the acquisition process of reviewed papers, outlining our taxonomy grounded in VR’s core elements and pivotal technological trajectories, and statistically analyzing the works within. Subsequently, we delve into the application of AI models, methodologies, and techniques across six research avenues: advanced AI-generated content representation, content rendering, content generation, physical simulation, virtual characters, and interaction, discussing their achievements. Concludingly, we summarize our findings, highlight existing challenges, and suggest potential avenues for future research.
The growing adoption of social virtual reality (VR) platforms underscores the importance of safeguarding personal VR space to maintain user privacy and security. Teleportation, a prevalent instantaneous locomotion method in VR, facilitates user engagement but can also inadvertently intrude upon personal VR space, thereby raising privacy concerns. This paper introduces three innovative negotiated teleportation techniques designed to secure user-to-user teleportation and protect personal space privacy, all under a unified small-group development framework. We have designed and evaluated three types of negotiated teleportation techniques: Sector technique for directional control, Distance technique for minimum social distance control, and Area technique for defining circular permissible teleportation areas. These techniques foster a collaborative approach to selecting teleportation points that respect personal space. To evaluate the efficacy of these techniques, we conducted a user study with 20 participants who performed social tasks within a virtual campus environment. The findings demonstrate that our techniques significantly enhance privacy protection and alleviate anxiety associated with unwanted proximity in social VR.
Editing 4D scenes reconstructed from monocular videos based on text prompts is a valuable yet challenging task with broad applications in content creation and virtual environments. The key difficulty lies in achieving semantically precise edits in localized regions of complex, dynamic scenes, while preserving the integrity of unedited content. To address this, we introduce Mono4DEditor, a novel framework for flexible and accurate text-driven 4D scene editing. Our method augments 3D Gaussians with quantized CLIP features to form a language-embedded dynamic representation, enabling efficient semantic querying of arbitrary spatial regions. We further propose a two-stage point-level localization strategy that first selects candidate Gaussians via CLIP similarity and then refines their spatial extent to improve accuracy. Finally, targeted edits are performed on localized regions using a diffusion-based video editing model, with flow and scribble guidance ensuring spatial fidelity and temporal coherence. Extensive experiments demonstrate that Mono4DEditor enables high-quality, text-driven edits across diverse scenes and object types, while preserving the appearance and geometry of unedited areas and surpassing prior approaches in both flexibility and visual fidelity.
The feature fusion of optical and Synthetic Aperture Radar (SAR) images is widely used for semantic segmentation of multimodal remote sensing images. It leverages information from two different sensors to enhance the analytical capabilities of land cover. However, the imaging characteristics of optical and SAR data are vastly different, and noise interference makes the fusion of multimodal data information challenging. Furthermore, in practical remote sensing applications, there are typically only a limited number of labeled samples available, with most pixels needing to be labeled. Semi-supervised learning has the potential to improve model performance in scenarios with limited labeled data. However, in remote sensing applications, the quality of pseudo-labels is frequently compromised, particularly in challenging regions such as blurred edges and areas with class confusion. This degradation in label quality can have a detrimental effect on the model’s overall performance. In this paper, we introduce the Difference-complementary Learning and Label Reassignment (DLLR) network for multimodal semi-supervised semantic segmentation of remote sensing images. Our proposed DLLR framework leverages asymmetric masking to create information discrepancies between the optical and SAR modalities, and employs a difference-guided complementary learning strategy to enable mutual learning. Subsequently, we introduce a multi-level label reassignment strategy, treating the label assignment problem as an optimal transport optimization task to allocate pixels to classes with higher precision for unlabeled pixels, thereby enhancing the quality of pseudo-label annotations. Finally, we introduce a multimodal consistency cross pseudo-supervision strategy to improve pseudo-label utilization. We evaluate our method on two multimodal remote sensing datasets, namely, the WHU-OPT-SAR and EErDS-OPT-SAR datasets. Experimental results demonstrate that our proposed DLLR model outperforms other relevant deep networks in terms of accuracy in multimodal semantic segmentation.
Shared virtual environments are becoming essential platforms for collaborative interaction and immersive entertainment, enabling users to be co-located and engage in activities together. The presence of surrounding virtual humans forms an environmental crowd, serving as a component of ambient stimuli in these environments. However, it remains unclear how crowd size affects users under different cognitive and motor demands. This study investigates the influence of crowd size on user performance, experience and social presence across three fundamental VR tasks: Spatial Locomotion, Memory Search, and Motor Coordination. We conducted a controlled within-subjects experiment, manipulating each task's crowd size at Small, Medium, and Large levels. Our results show that crowd size significantly impacts user performance, experience, and social presence, but these effects are task-dependent. While Medium size can enhance performance, Large size in cognitively demanding tasks may induce attentional blindness and diminish sensitivity to social cues. Task functionality further shapes how users perceive and respond to virtual crowds. Additionally, users' preferences for crowd size varied across different tasks, and most participants expressed a desire for control over the number of visible avatars. These findings provide novel insights into human crowd perception mechanisms, revealing cross-task perceptual variations that pave the way for further exploring crowd perception in shared virtual environments.
Existing radiance field-based head avatar methods have mostly relied on pre-computed explicit priors (e.g., mesh, point) or neural implicit representations, making it challenging to achieve high fidelity with both computational efficiency and low memory consumption. To overcome this, we present GPAvatar, a novel and efficient Gaussian splatting-based method for reconstructing high-fidelity dynamic 3D head avatars from monocular videos. We extend Gaussians in 3D space to a high-dimensional embedding space encompassing Gaussian’s spatial position and avatar expression, enabling the representation of the head avatar with arbitrary pose and expression. To enable splatting-based rasterization, a linear transformation is learned to project each high-dimensional Gaussian back to the 3D space, which is sufficient to capture expression variations instead of using complex neural networks. Furthermore, we propose an adaptive densification strategy that dynamically allocates Gaussians to regions with high expression variance, improving the facial detail representation. Experimental results on three datasets show that our method outperforms existing state-of-the-art methods in rendering quality and speed while reducing memory usage in training and rendering.
Motion retargeting is an active research area in computer graphics and animation, allowing for the transfer of motion from one character to another, thereby creating diverse animated character data. While this technology has numerous applications in animation, games, and movies, current methods often produce unnatural or semantically inconsistent motion when applied to characters with different shapes or joint counts. This is primarily due to a lack of consideration for the geometric and spatial relationships between the body parts of the source and target characters. To tackle this challenge, we introduce a novel spatially-preserving Skinned Motion Retargeting Network (SMRNet) capable of handling motion retargeting for characters with varying shapes and skeletal structures while maintaining semantic consistency. By learning a hybrid representation of the character's skeleton and shape in a rest pose, SMRNet transfers the rotation and root joint position of the source character's motion to the target character through embedded rest pose feature alignment. Additionally, it incorporates a differentiable loss function to further preserve the spatial consistency of body parts between the source and target. Comprehensive quantitative and qualitative evaluations demonstrate the superiority of our approach over existing alternatives, particularly in preserving spatial relationships more effectively.
In social virtual reality (VR), maintaining appropriate interpersonal distance is essential for user comfort and privacy. However, most existing locomotion methods provide limited support for respecting personal space, leaving users vulnerable to unintentional or socially inappropriate intrusions. To address this issue, we propose potential field-guided teleportation, a proactive locomotion framework consisting of two method implementations that dynamically adjust teleportation targets based on real-time interpersonal proximity, preventing entry into others' personal spaces without explicit user intervention. We evaluate our technique through two user studies: a preliminary study exploring energy-based constraint parameters, followed by a comparative study against conventional and negotiated teleportation methods. Experiments were conducted in socially interactive VR scenarios populated with simulated users exhibiting human-like behaviors. Results demonstrate that our methods reduce perceived social anxiety while maintaining locomotion efficiency and usability. This work presents a socially-aware locomotion strategy that balances personal space protection with effective and socially appropriate movement in shared virtual environments.
The locomotion and interaction of multi-user groups are critical components of social virtual reality (VR), where users collaboratively navigate shared spaces defined by their relationships. As metaverse and social VR platforms evolve, safeguarding group spatial integrity becomes paramount. While teleportation-widely adopted for efficient navigation-enhances communication, it risks unintended intrusion into group spaces, compromising privacy and security. This paper presents two novel negotiated user-to-group teleportation techniques, paired with dynamic zone computation methods, to address these challenges. Our approach enables guest users to join groups through spatially aware teleportation, mediated by real-time negotiation. The negotiation interaction process is designed to facilitate users to negotiate teleportation locations efficiently and smoothly. To validate our techniques, we conducted a user study with 36 participants in a VR art museum environment, where they performed collaborative social-tour tasks. The findings demonstrate that our techniques significantly enhance group privacy protection, effectively support user-to-group negotiation of teleportation joining requirements, and alleviate anxiety associated with unwanted proximity within social VR groups.
We present Neural 3D Strokes, a novel technique to generate stylized images of a 3D scene at arbitrary novel views from multi-view 2D images. Different from existing methods which apply stylization to trained neural radiance fields at the voxel level, our approach draws inspiration from image-to-painting methods, simulating the progressive painting process of human artwork with vector strokes. We develop a palette of stylized 3D strokes from basic primitives and splines, and consider the 3D scene stylization task as a multi-view reconstruction process based on these 3D stroke primitives. Instead of directly searching for the parameters of these 3D strokes, which would be too costly, we introduce a differentiable renderer that allows optimizing stroke parameters using gradient descent, and propose a training scheme to alleviate the vanishing gradient issue. The extensive evaluation demonstrates that our approach effectively synthesizes 3D scenes with significant geometric and aesthetic stylization while maintaining a consistent appearance across different views. Our method can be further integrated with style loss and image-text contrastive models to extend its applications, including color transfer and text-driven 3D scene drawing. Results and code are available at http://buaavrcg.github.io/Neural3DStrokes.