
ABSTRACT Emergency procedures are an important domain for safety education. Modern multimedia technologies aim to aid this process by replacing real‐life drills and hazard procedures with virtual simulations. However, there is a limited amount of existing research on this topic, especially focusing on the Western Balkans region. This paper presents a 3D virtual learning environment for emergency procedures in the form of a serious game. The virtual user manual uses the 2D fire evacuation plan of a university building in Ohrid, North Macedonia, for teaching civilians about multiple emergency procedures emergency exits, fire alarms, fire extinguishers, and wall hydrants. A case study was performed on 90 participants from Bosnia and Herzegovina and North Macedonia. Four different scenarios of interest were utilized, general knowledge about emergency procedures, emergency exit simulation, virtual user manual simulation, and emergency procedures involving a cultural heritage site. The achieved results indicate that participants are open to the usage of modern multimedia technologies, that familiarity with the presented location affects the success of learning, and that there is a strong potential for successful usage of virtual user manuals in the region of the Western Balkans in the future.
ABSTRACT In this article, we present a novel, streamlined simulation method for visualizing the dynamics of spinning tops, specifically focusing on convex shapes with explicit representations. Our simulation approach contributes to the understanding of the instricate dynamics of the Gömböc by carefully building an accurate underlying mathematical model. This result, alongside the capability to model classic spinning top types, namely the Ellipsoid, Tippe Top, Euler's disk, and Rattleback, emphazises the versatility of our method. Designed for ease of implementation and understanding, our approach allows for straightforward adaptation to a wide range of such shapes, making it a valuable tool for researchers, educators, and students investigating the world of spinning tops, particularly in settings with limited computational resources. Lightweight design ensures efficient execution on CPUs with single‐thread processing capabilities, thus enhance educational exploration and dynamic understanding of complex rotational motion within web‐based environments.
ABSTRACT Holistic talking human animation aims to synthesize coordinated facial expressions and body poses that are consistent with spoken content and human communication patterns. This task requires both realistic motion generation and accurate temporal alignment between speech and motion, which remain challenging for existing methods that model facial and bodily motion separately or operate in high‐dimensional motion spaces. We present HoloDiff, a latent diffusion framework for holistic talking human animation. Given an input speech signal, HoloDiff encodes facial and bodily dynamics into a compact and structured latent representation that preserves anatomical structure and cross‐part interactions. A speech‐conditioned diffusion process is then applied in this latent space to model long‐range temporal dynamics and generate motion sequences aligned with audio input. By performing diffusion in a structured latent space, HoloDiff reduces modeling complexity while improving temporal coherence and audio–motion alignment. Extensive experiments on the SHOW dataset demonstrate that HoloDiff produces more realistic and expressive talking human animation and achieves better multimodal correspondence between speech and motion compared to baselines.
ABSTRACT Virtual Reality (VR) has gained significant momentum in creative endeavors due to its unique characteristics of immersion, interaction, and imagination. However, effectively harnessing its “imagination” potential—specifically how VR content influences creative ability and empowers humans in goal‐oriented tasks like product design—remains a challenge. This paper presents a tailored Immersive Ideation System (IIS) for product appearance design. The system enables designers to perform sketching, modeling, and conceptual review within integrated VR scenes, fostering a fluid and seamless ideation workflow. To evaluate the efficacy of the IIS, we conducted a comprehensive task‐centered user study involving 24 industrial design students and 6 professional designers. Our analysis focuses on the influence of VR environments on creative behaviors and design outputs. The results suggest that while specific VR conditions relate to differences in design flexibility and originality, they also introduce potential trade‐offs with design fixation, offering tentative insights for the configuration of future immersive design workspaces.
This paper introduces a reinforcement learning framework for simulating ego-motorcycle behavior in mixed traffic for animation-oriented applications. A local path planner uses ray-based perception to guide navigation, while a tailored reward structure promotes safe interactions and penalizes hazardous maneuvers. To support action selection, the planner also constructs a risk-probability map that captures nearby vehicles' speeds and distances. Curriculum learning gradually increases traffic complexity during training, enabling the policy to adapt to a wide range of scenarios. Experimental results show that the trained policy produces collision-free trajectories, responsive deceleration, and well-regulated maneuvering. It also occasionally exhibits assertive or borderline-risky interactions that enhance behavioral naturalness. These findings demonstrate that the framework can generate structured yet varied motorcycle motion. Potential applications include entertainment media, educational visualizations, and driver-training simulations, enriching animated content for learning and engagement.
Forests are crucial terrestrial ecosystems. To understand long-term forest community evolution driven by multiple environmental factors, we constructed a 3D forest stand spatiotemporal evolution framework featuring synchronous bidirectional coupling between terrain, hydrology, radiation, and vegetation. First, we simulated topographical evolution using a physics-based procedural erosion method and introduced a multi-layer soil moisture model and individual tree growth response mechanisms to reflect-topography interactions, thereby establishing a realistic environmental basis for forest stand evolution. Second, leveraging real environmental data, we simulated growth responses and biomass changes of forest stands under different precipitation and radiation conditions, elucidated response mechanisms of individual tree attributes to environmental changes, and achieved intuitive evolution of 3D forest stands. The framework advances beyond unidirectional environmental forcing models by integrating hydraulic erosion, soil moisture dynamics, and slope-aware radiation within a unified monthly timestep, enabling co-evolutionary simulation of forest stands under dynamic landscapes. Finally, the computer-based model incorporated natural disaster events such as fires and droughts, with real-time interaction and visualization capabilities, supporting immersive and responsive forest landscape simulation and detailed spatiotemporal evolution analysis, enabling assessment of the dynamic recovery processes of forest stands under multiple disturbance scenarios.
Sand painting is a visually distinctive art form characterized by granular textures and diverse expressive techniques, where sand grains are manipulated through various hand movements such as waving, seeping, sweeping, and stroking. However, traditional stylized methods often fail to capture the fine textures and diverse techniques unique to sand painting. In this paper, we propose a Feature-Oriented Sand Painting (FOSP) inspired by real sand painting techniques, aiming to produce sand paintings with authentic sandy textures and diverse techniques. Our FOSP comprises three main modules: overall sand waving generation, extraction and drawing of sand reduction regions, and simulation of sand grain accumulation with multiple techniques. The first module generates fine sand-grain textures, whereas the latter two focus on contour rendering to enrich detail expression. Experiments validate that our FOSP can generate high-quality sand paintings rapidly, outperforming existing methods in sand painting generation tasks.
ABSTRACT The rapid convergence of virtual reality (VR) and artificial intelligence (AI) has resulted in a redefined ethical framework for digital interaction. As AI‐driven avatars gain cognitive and emotional autonomy, they transition from passive representations to active participants capable of shaping moral reasoning, social behavior, and value construction. The present study explores the ethical transformations that are emerging from this integration, focusing on how human‐machine interaction reconfigures traditional concepts of responsibility, agency, and moral judgment. By employing a systematic and thematic analysis of the extant literature, the research identifies key trends such as the shift from user‐centred ethics towards distributed and relational ethical systems embedded within technological frameworks. The findings indicate that immersive environments exert a significant influence on users' empathy and moral sensitivity, concomitantly giving rise to novel dilemmas pertaining to authenticity, consent, and algorithmic accountability. The study posits that the evolution from human–avatar symbiosis to AI–avatar autonomy signifies a fundamental transformation in the nature of ethical engagement within virtual ecosystems. These insights emphasize the necessity for inclusive, transparent, and value‐sensitive approaches to guide the design and governance of future AI‐integrated VR environments.
Current BIM visualization systems tessellate parametric geometry into triangle meshes before rendering, which increases memory usage, introduces faceting on curved surfaces, and adds significant export-time overhead. We present an end-to-end system that preserves parametric shape definitions from BIM models through to GPU ray tracing, bypassing mesh conversion for supported parametric shapes. Our exporter extracts parametric shapes from a BIM authoring tool and encodes them into a compact binary format with two-tier caching to avoid redundant extraction. Our renderer maps preserved BIM primitives to an instanced GPU scene layout and performs direct ray-primitive intersections, with mixed-precision parameter encoding and ray advancement improving memory efficiency and numerical robustness. Evaluation on architectural projects shows that our system achieves a lower GPU memory footprint, produces mathematically exact curved surfaces free of tessellation artifacts, and exports models faster than standard workflows.
Point cloud saliency prediction, which models Human Visual System (HVS) attention, is essential for applications like Computer Animation and eXtended Reality (XR). However, most existing research focuses on static scenes, and extending these models to dynamic point clouds is hindered by motion-induced attention shifts and a lack of ground-truth dynamic datasets. To address this, we propose the Weakly-Supervised Dynamic Point Cloud Saliency Prediction Network (WDPS), one of the first frameworks to utilize classification supervision for dynamic 3D saliency. WDPS integrates saliency-oriented feature mining, multiscale semantic-detail fusion, and a temporal smoothness constraint to generate stable dynamic saliency maps. Extensive experiments demonstrate that our method is preferred by observers in over 78% of comparisons against static baselines. Furthermore, WDPS achieves superior Velocity-Weighted Saliency (VWS) scores on highly dynamic sequences and significantly enhances downstream tasks performance.
Music-driven motion generation has attracted increasing attention, yet conducting remains underexplored despite its central role in musical communication. We address this gap with a focus on choral conducting. We introduce Maestro3D, a large-scale 3D dataset of professional performances, and propose a beat-aware diffusion framework enhanced by a novel phase-based beat representation that explicitly encodes rhythmic structure. Both quantitative and qualitative evaluations show that our approach achieves superior accuracy, realism, and synchronization compared to existing methods.
This paper presents a comparative study of two variants of the user interface in a virtual reality (VR) museum application: an interface designed in accordance with the principles of universal design (UD) and an interface that does not meet these principles (non-UD). The application was developed in the Unity engine and run on Meta Quest Pro headsets with built-in eye tracking. During the completion of a set of tasks, eye-tracking data (including heatmaps and gaze scanpaths) and task completion times were recorded, and subjective ratings were then collected using the standardized system usability scale (SUS). The results indicate an advantage of the UD variant over the non-UD variant: the mean SUS score was higher for UD (73.75) than for non-UD (65.31). A Wilcoxon signed-rank test confirmed the statistical significance of this difference (Z = 4.87, p < 0.001, r = 0.49). Moreover, tasks were completed faster in the UD variant (total time 40.41 s) than in the non-UD variant (57.85 s). This difference was also statistically significant (Z = 5.61, p < 0.001, r = 0.56). Eye-tracking observations supported these differences, suggesting more efficient interface exploration in the UD variant.
This work presents a comparative study evaluating how recent foundation models, capable of handling various computer vision tasks, perform at simulating a virtual agent's perception and using that perception to navigate the environment and achieve its goals. To achieve this, we selected two recent foundation models well known in the literature and compared them with a recent procedural approach by replicating these models' evaluation methods. We employed these models for evacuating a burning building, aiming to identify either exit signs or people to follow in the environment to find an escape route. Our results show that out-of-the-box foundation models were unable to outperform a specialized procedural method and performed significantly worse on the task, although they still exhibit strong adaptability to previously unknown contexts. However, among the two foundation models tested, we found that a VLM capable of image reasoning and detection performed significantly better than a self-supervised generalized computer vision model.
We present a monocular camera calibration method, dubbed CameraVQ. It reformulates monocular camera calibration as classification over vector-quantized camera intrinsics. Existing methods based on geometric cues or direct regression often suffer from unstable optimization and poor generalization. In contrast, CameraVQ learns a discrete codebook of camera intrinsics and predicts the latent code from a single image. This discrete formulation constrains predictions to a statistically learned manifold of valid configurations, enabling robust calibration and strong generalization. Through extensive evaluation on diverse calibration benchmarks, CameraVQ achieves state-of-the-art performance on the majority of datasets.
Visually induced motion sickness (VIMS) is one of the major obstacles to the broader adoption of virtual reality (VR) technology. As visual-vestibular conflict is considered a major factor contributing to VIMS, real-time prediction from visual cues remains crucial for timely intervention and for improving user experience. In this paper, we present a real-time VIMS prediction model that utilizes a hybrid architecture of a 3D Convolutional Neural Network and a Long Short-Term Memory network to capture both short and long-term features. We used low-resolution frames and replaced optical flow with frame difference maps and lateral/vertical displacement components to reduce data complexity and accelerate model inference. Experimental results on two datasets demonstrate that our model outperforms existing vision-based approaches, such as VR-SP and VR-SA, with an RMSE of 3.77 and a PLCC of 0.935. Importantly, the average processing time for each 108-frame video is less than 0.83 s. In addition, the method's real-time VIMS prediction capability has been verified in a Unity3D-based VR environment, with an end-to-end latency of 192 ms. These results highlight the model's advantages in terms of both efficiency and accuracy, making it a promising solution for VIMS-aware VR applications.
Queue cutting is a frustrating experience in the real-world when a person is waiting to receive service as it violates the normative behavior of first-in-first-out ordering of queues. As social experiences move to virtual worlds users may experience human and non-playable characters (NPCs) attempting to cut the queue, thereby adding to the negative experiences of waiting. While normative behavior for activities such as joining an existing conversational group or walking between agents engaged in conversation have been studied in virtual reality, perceptions toward queue cutters have not been studied. We conduct a study on understanding how users perceive a queue-cutting NPC in a virtual doctor's office reception area. During the wait, a queue-cutting NPC requests an in-queue NPC to cut in ahead of them and is either allowed or denied entry into the queue. Using data on subjective responses from 45 participants we find significant differences in perception of time, frustration, and likelihood of exit when the queue-cutting NPC is allowed in. Our work enables future research on detecting situations capable of generating user frustration, and providing appropriate intervention via the VR environment, mitigating negative experiences, and ensuring timely service.
Realistic simulation of pedestrian movement in natural environments remains challenging due to the limited interaction between agents and dynamically changing terrain. In real-world settings, pedestrian trails emerge through repeated footstep interaction with deformable terrain such as snow, grass, or soil, and these trails subsequently influence navigation behavior and crowd flow. However, existing crowd simulation approaches typically treat terrain as static or rely on predefined paths, limiting both visual realism and behavior consistency. We present an environment-aware crowd navigation model that integrates footstep-driven trail formation, directional trail encoding, and trail-aware pathfinding to produce emergent, visually grounded pedestrian behavior. Our approach models terrain as a deformable surface represented by a dynamic height and trail-intensity field updated at each footstep. As agents traverse the environment, footsteps increment local trail intensity and modify terrain appearance, producing persistent visual trails. These trails are simultaneously encoded into a 2D grid used for navigation. Pathfinding costs are dynamically adjusted using accumulated foot traffic, encouraging agents to follow existing trails while still allowing divergence when beneficial. Results show improved visual plausibility, natural trail reuse, and coherent crowd flow compared to static-terrain baselines.
City-scale 3D urban generation requires planning-level semantic grounding from user intent and scalable geometric synthesis with structural validity and editability. Procedural content generation (PCG) offers controllability and scalability, but is hard to author due to high-dimensional parameters and nonintuitive workflows. Meanwhile, directly generating city geometry or scripts from text with LLMs can suffer from weak large-scale consistency and limited geometric validity, hindering downstream editing and engine deployment. We present Text-to-3D City, a plan-then-execute framework that couples an LLM-based City Planner with a PCG-based Implementer. Given a natural language description, the Planner grounds textual intent into a structured city plan by composing PCG parameters via a schema and in-context exemplars. The Implementer deterministically executes road generation, block extraction, lot subdivision, and asset placement with validity checks and reproducible seeding to synthesize an engine-ready 3D city. Experiments on multi-view renderings evaluate text-scene alignment, diversity, realism, and runtime, demonstrating rapid generation and scalability to large urban scenes.
Text-driven 3D human motion editing aims to modify an existing motion sequence following natural language instructions, which is a crucial task for character animation, virtual agents, and motion authoring. Recent diffusion-based methods have shown remarkable success in text-to-motion generation. Editing existing motions requires precise spatiotemporal control to localize modifications while preserving context. Current diffusion-based motion editing methods lack explicit fine-grained control over when and how strongly to edit. To address this, we propose TM-Edit, a text guided Diffusion-Transformer based motion editing framework which introduces learned temporal soft masks to provide explicit frame-wise editing guidance. The proposed model predicts an editing intensity mask to encode high-level intent from both the source motion and the text instruction. This mask is then used to modulate source motion features within a conditional diffusion process via an uncertainty-aware gating mechanism, ensuring robust training and inference. Additionally, a feature semantic alignment loss is employed by using a pre-trained motion retrieval model to enhance cross-modal consistency. Extensive experiments on the MotionFix benchmark dataset demonstrate that our approach achieves state-of-the-art performance. Code will be made publicly available.
Traditional physics-based fluid simulations typically rely on manual modeling and incremental adjustments to achieve desired effects, which can limit objectivity and generalizability to new scenarios. To address these challenges, we propose a novel neural fluid simulator that integrates visual priors from 2D image sequences with physically constrained continuous convolution. Specifically, we extract and refine point clouds from image sequences, then infer the kinetic properties of the fluid. We introduce an energy-based physical constraint and incorporate it into a continuous convolution solver. By iteratively optimizing these inputs to enforce physical laws-particularly incompressibility-the solver produces accurate fluid motion predictions. Our approach uniquely combines visual data and physical constraints, enhancing the realism and accuracy while providing stronger generalization of fluid simulations.