
The arrangement of objects in a 3D gaming environment is key to guiding players' visual attention. The human visual system is naturally drawn to salient areas, characterized by high contrast and importance, which directly affect performance in games like hidden-object and first-person shooters. This study investigates how improving user experience can be achieved by varying game difficulty according to objects' saliency. We created two games with three distinct levels of difficulty in which players utilized virtual reality to locate targets in various environments. Game difficulty was modulated by the saliency of these objects within their environments. In the second experiment, we applied the same structure but focused on how varying the materials and colors of objects influenced detection. Our findings show that controlling object saliency and visual characteristics can effectively manage game difficulty, offering new possibilities for adaptive and immersive gaming experiences.
The vision of Industry 5.0, which addresses humanistic, adaptable and sustainable method, has evolved rapidly, particularly in digital twin-driven (DT) utilization in industry, for example, BMW factory and Tata steel. Nevertheless, it is no secret that the development of DT-powered virtual reality (VR) environments is hard work, which denotes there is untapped potential. Thus, a reusable DT model is required to speed up the creation and scaling of DT-driven VR environments. The objective of this study is for seeking solutions to make DT-driven VR contexts faster in an industrial environment. In particular, this study presents a way to develop the settings more effectively. The highlight of this study, and related industrial instances, is the development of safety education for human-robot collaboration in manufacturing contexts utilizing DT-driven VR settings. A formwork mode is introduced, which involves specific formworks and the entire framework for disposing the formworks in order to more quickly amend DT-driven VR environments to satisfy the requirements of a particular instance. The formwork is established applying two different industrial instances from manufacturing scenario.
Emergency plan training is very important for subway staff to respond quickly in the event of a fire emergency in a subway station. Compared to traditional training methods, immersive virtual reality serious games (IVR SGs) are gradually attracting attention. However, there are few studies currently on the application of this training method to fire emergency plan training for subway staff. Therefore, an IVR SG for subway intern staff fire emergency plan training was designed and developed. Firstly, the prototype development process for this IVR SG was described, then the effectiveness of the developed IVR SG training was validated through being compared with traditional written text training from the perspectives of knowledge acquisition and retention, self-efficacy improvement, and users' virtual reality (VR) experience. The results showed that the developed IVR SG was effective in terms of knowledge acquisition and self-efficacy improvement, and was more effective than the traditional written text training in short-term and long-term training effectiveness. In addition, the VR experience that the developed IVR SG provided to users was also acceptable. The study can provide some insights for the prototype development and validation of effectiveness on IVR SG for the emergency plan training.
Although numerous studies have explored relaxation and sleep aid through Autonomous Sensory Meridian Response (ASMR) videos or conventional Virtual Reality (VR) relaxation methods, the integration of VR 3D animation with ASMR and its comparison to traditional VR relaxation methods remains underexplored. To address this gap, we investigate a standardized process for creating a VR-based ASMR animation game and its impact on triggering the ASMR tingling sensation in VR environments. We also developed a VR 3D environment game featuring four natural environments, along with one ASMR video as a control group. A comprehensive experiment was conducted to compare the effectiveness of these three relaxation methods. Forty-seven participants aged 18-35 from Bournemouth University were recruited and divided into three experimental groups. Participants' emotional and physiological responses were monitored using both subjective questionnaires and physiological data collection that is, heart rate (HR) and electrodermal activity (EDA). Our findings show that VR-based ASMR animation game effectively triggers the ASMR tingling experience and offers superior relaxation, sleep assistance, and emotional regulation compared to watching ASMR videos and conventional VR relaxation methods, resulting in a significant reduction in anxiety and stress, as well as increased feelings of calmness and sleepiness.
Crowd anomaly detection is a critical aspect of ensuring public safety in various domains such as surveillance and security. Ensuring public safety in crowded environments requires accurate and efficient crowd anomaly detection. This research proposes an innovative approach to crowd anomaly detection using the SqueezeNet-ImpLinknet architecture. The input images are first preprocessed using a median filtering technique. Then, object segmentation takes place using an Improved Mask Region-based CNN. It incorporates batch normalization, ReLU activation, and an advanced Scale Dot Product attention mechanism to improve segmentation accuracy and computational efficiency. Subsequently, features such as the Improved SLBT feature, capturing shape and texture information, color features, and LGTrP features are extracted. Then, anomaly detection is performed using a hybrid model that integrates SqueezeNet and Improved Linknet models. The Improved LinkNet model enhances feature representation by integrating an attention mechanism in the encoder and a novel ReLUSignmax activation function in the decoder, overcoming limitations of conventional architectures. The approach is evaluated on the widely used UCSD Anomaly Detection Dataset, achieving superior performance with accuracy ranging from 0.939 to 0.975 and a specificity of 0.987 at 90% training data. The proposed approach offers a robust solution for intelligent surveillance in crowded environments.
This study proposes a low-cost, Arduino-based walking interface designed to enhance user immersion in virtual reality (VR) content. The interface detects various walking motions such as walking, running, and limping, and reflects these movements in the virtual environment to deepen the user's sense of immersion. To achieve this, a sensor device based on Arduino was developed to analyze acceleration data from a gyroscope sensor. This data was then integrated with the Unity3D engine to synchronize character animations and first-person visual effects. Additionally, a multi-user environment was implemented using Bluetooth connectivity with mobile devices, allowing both head-mounted display (HMD) and non-HMD users to share the same virtual experience. Owing to its low cost and high compatibility, this interface can provide an immersive VR experience not only for general users but also for those with mobility impairments, suggesting its potential for applications in fields such as healthcare, education, and psychological rehabilitation.
This paper presents a prototype XR system for disaster medical response, demonstrating the feasibility of real-time interactive elastodynamics simulations in emergency scenarios. The system delivers an end-to-end workflow utilizing XR technology, from on-site data acquisition to remote simulation. Specifically, we propose an image-guided mesh-processing pipeline that converts photographs of injured individuals into solver-ready tetrahedral meshes. We also develop a constraint-based elastodynamics solver capable of simulating deformable bodies and visualizing internal stresses. Additionally, the system integrates multiple advanced XR devices and addresses the coordinate-alignment problem between these devices and the simulator. We validate the system's performance in both AR/VR modes, under textured and stress-visualization configurations, and demonstrate its applicability for remote medical guidance. Beyond whole-body elastic simulations, we conduct preliminary organ-level experiments to inform future remote surgical applications. This prototype, validated using a two-room setup, provides a feasible solution for remote emergency medical response.
Effective audiovisual cueing can significantly enhance learners' attention to educational resources in the Virtual Reality (VR). However, predicting the impact of multimodal cueing on learners' attention in immersive teaching environments remains a challenging task. To address this, we propose a deep learning model named Attention Prediction Model (APM). This model employs RevFCN to extract visual and auditory cue features and incorporates a tailored Upsample-Aggregation Fusion Module (UAFM) to integrate multimodal representations. Additionally, an SANet is introduced to effectively combine the advantages of spatial and channel attention. Trained on our constructed dataset, APM achieved an attention prediction accuracy of 81.6%. These findings offer both theoretical and practical implications for the application of multimodal cueing in VR-based instructional design.
Interactive computer technology is deeply integrated into traditional teaching methods. The traditional teaching of physics experiments in secondary schools suffers from the inability of teachers to provide timely guidance to students, the difficulty of controlling experimental variables, and the lack of uniformity in evaluation criteria. To address the aforementioned issues, we have developed an innovative system to improve secondary school physics education using computer vision-based interaction with virtual humans and sensors. The proposed system captures experimental data in real time so that student performance can be accurately monitored and assessed. Teachers can effortlessly configure experiments through simple coding, while the system leverages a multimodal macrolanguage model to offer contextual feedback and guidance. The system generates a virtual teacher that offers step-by-step guidance and real-time feedback. Usability tests indicate that the system significantly improves student engagement and comprehension of complex physics concepts, highlighting its potential to transform traditional science education. The advantage of real-time assessment and guidance in secondary school physics experiments is that it enables students to grasp abstract concepts in a more intuitive and comprehensible manner.
Comprehensive observation of target area with pedestrians through pan-tilt-zoom (PTZ) camera networks is crucial in various surveillance applications. However, the dynamic configuration of PTZ cameras increases the difficulty of coordinating multiple cameras to monitor large-scale scenes. Since coverage control in PTZ camera networks has been proven to be an NP-hard problem, many studies have adopted virtual potential field (VPF) algorithms to efficiently obtain approximate solutions. The VPF methods treat camera viewpoints as charged particles. Through repulsive forces between these particles, PTZ camera networks expand scene coverage and reduce overlap between camera fields of view (FoVs). However, VPF-based methods cannot leverage the scene layout and target priority information, failing to cover pedestrians and other critical areas. In this work, we introduce a unified perception quality measure framework that quantifies surveillance importance for scenes, cameras, and pedestrians. Building on this framework, we design a coverage control scheme using a perceptive quality-based virtual potential field. This scheme models target regions and pedestrian priorities as virtual gravitational and attractive forces. It maximizes coverage of key regions, minimizes camera overlap, and supports high-resolution monitoring and tracking of pedestrians. Extensive experiments show that our approach outperforms state-of-the-art methods, achieving superior scene and pedestrian coverage performance.
Dance generation is a significant research area in computer arts and artificial intelligence. This study proposes a novel framework to enhance dance controllability and personalization through multimodal and multi-granularity control. The framework establishes global choreographic control of long sequences via music and dance style factors, while accommodating local style variations. Simultaneously, it enables fine-grained local control using style, text, and temporal factors for motion refinement. We develop two cross-modal Transformers: the LS-M2D model merges music and dance style features for local style-controllable dance generation, and the LT-SM2D model integrates textual guidance with music and dance style features for time-constrained local control. Experimental results demonstrate enhanced motion quality, effective multi-granularity style control, and precise text-guided flexibility. This provides valuable technical support for personalized intelligent dance generation systems.
Advances in microelectronic components and high-speed networks have enabled the widespread application of virtual reality (VR) technology in education. However, insufficient attention to knowledge visualization in VR learning has resulted in disorganized knowledge structures, comprehension difficulties, and mismatches between user experience and learning achievement. Therefore, we propose a Knowledge Cube (KC)-based visualization method to standardize knowledge encoding in VR learning. During courseware development, the instructor defines discrete knowledge as Events, organizes them into Event Groups, and populates data into a KC model to generate VR courseware. In subsequent VR learning, when learners search for task-relevant knowledge using the provided retrieval method, the KC model presents the corresponding events within interactive scenarios according to its predefined structure. Comparative experiments on different knowledge visualization methods revealed that, in VR learning, the KC method outperforms other VR approaches in both learning performance and efficiency. This method effectively guided learners to focus on the learning content and optimized the knowledge encoding in VR learning. This provides an operational framework for knowledge encoding in VR courseware design and emphasizes the importance of supporting effective learning behaviors over merely pursuing immersion, presenting a new perspective for refining the design approach of VR courseware.
Generating diverse and realistic movements has long been a central challenge in computer graphics. Generative Adversarial Networks (GANs) remain a compelling solution due to their ability to perform well even with limited training data. However, traditional GANs generate samples directly, which can lead to the omission of certain data patterns. To address this limitation, we introduce SinMDGan , a hybrid deep learning framework for single-motion synthesis that leverages a Diffusion-GAN model. Our approach integrates the strengths of GANs, which capture global motion characteristics, with diffusion techniques, which refine local details, ensuring both authenticity and diversity in generated movements. Unlike conventional cascaded GANs, our framework employs a single generator-discriminator pair, utilizing different diffusion time steps to synthesize novel and diverse motions from a single short sequence. Experimental evaluations demonstrate the effectiveness of our model in achieving stable data distribution coverage and enhancing output diversity. Additionally, we showcase various applications, including motion composition and long-sequence generation, highlighting the versatility of our approach.
Sign language recognition (SLR) requires interpreting dynamic hand gestures with complex variations in shape, orientation, motion, and spatial configuration. Conventional models such as U-Net and ResNet offer strengths in segmentation and feature extraction, respectively, but face critical limitations. U-Net struggles with retaining fine spatial details in cluttered backgrounds and lacks temporal modeling, while ResNet can lose motion continuity and suffers from vanishing gradient issues in deeper architectures. To overcome these challenges, we propose the holistic sign language interpretation network (HSLIN), a novel deep learning framework tailored for Indian sign language (ISL) recognition. HSLIN incorporates three key innovations: Uniformed frame isolation and augmentation (UFIA) for standardized preprocessing and noise removal, synaptic gesture movement analysis (SGMA) for capturing detailed motion using keypoint detection and optical flow, and a hybrid architecture combining U-Net-based segmentation with an enhanced ResNet-TC50V2 backbone. The novelty lies in fusing spatial precision with deep temporal modeling through bottleneck layers and temporal convolutional layers (TCL), enabling the model to effectively learn gesture patterns over time. Experimental results on the ISL-CSLTR dataset demonstrate that the proposed method achieves an accuracy of 99.9%, a precision of 100%, recall of 99.9%, and an F1-score of 100% across 14 word-level sign classes. Furthermore, an ablation study confirms the critical role of each architectural component in achieving optimal performance. These outcomes clearly establish the robustness, efficiency, and uniqueness of the proposed HSLIN framework, positioning it as a powerful solution for real-world ISL recognition and communication accessibility for the deaf and hard-of-hearing community.
Recent advances in Score Distillation Sampling (SDS) have enabled text-driven 3D human generation, yet the standard classifier-free guidance (CFG) framework struggles with semantic misalignment and texture oversaturation due to limited model capacity. We propose a novel framework that decouples conditional and unconditional guidance via a dual-model strategy: A pretrained diffusion model ensures geometric stability, while a preference-tuned latent reward model enhances semantic fidelity. To further refine noise estimation, we introduce a lightweight U-shaped Swin Transformer (U-Swin) that regularizes predicted noise against the reward model, reducing gradient bias and local artifacts. Additionally, we design a time-varying noise weighting mechanism to dynamically balance the two guidance signals during denoising, improving stability and texture realism. Extensive experiments show that our method significantly improves alignment with textual descriptions, enhances texture details, and outperforms state-of-the-art baselines in both visual quality and semantic consistency.
This study explores how environmental design elements in library spaces influence human psychophysiological responses using virtual reality (VR). Thirty participants experienced VR simulations of library reading areas, with variations in wall color, flooring material, and lighting intensity, whereas electroencephalography (EEG) and galvanic skin response (GSR) measured physiological reactions alongside subjective ratings. Moderate lighting (20,000-30,000 cd) minimized arousal and supported attention, while white walls enhanced relaxation via increased alpha brain activity. Green plant walls slightly boosted attention-related beta activity, and wood flooring was rated highest for comfort and naturalness. VR enabled precise control of design variables, advancing environmental psychology research. These findings offer evidence-based guidelines for designing public spaces like libraries to enhance user well-being and cognitive performance, with implications for educational and public buildings.
We propose new crowd simulation methods to virtually reenact the Itaewon disaster that occurred in 2022 due to the extreme crowd density. Conventional techniques make it challenging to simulate diverse, extremely dense crowd behaviors such as crowd surge, fluidization, and falls observed at the Itaewon disaster. This paper proposes a kinodynamic agent simulation combining kinematic agents for low-density crowds, hydrodynamic and hydrostatic agents for high-density crowds, and articulated passive agents for high-density crowds in dense contact. In order to perform co-simulation among heterogeneous agent types, we use a message-passing mechanism to share relevant kinematic and dynamic information among agents and make agent-type transitions based on crowd density and contact forces. Experiments show that the proposed hybrid simulation approach can accurately reenact crowd phenomena observed at the Itaewon compared to the CCTV footage. Moreover, our ablation study supports the use of kinodynamic agents to faithfully reconstruct the Itaewon crowd behavior. Furthermore, we run three what-if scenarios to explore the possibilities of using our techniques to help prevent incidents in the future. Finally, to demonstrate the applicability of our proposed methods to other types of extreme crowd behaviors besides the Itaewon disaster, we simulate two other real-world crowd incidents using our techniques.
The proposed MotionBlend GAN model marks a significant step forward in video synthesis by blending the motion from a source video with the appearance of a target person's image. As training progresses, the model improves video creation by enhancing the smoothness and natural flow of motion, resulting in more coherent and lifelike videos. Using advanced techniques like MoBConv blocks of EfficientNet-B7, OpenPose for precise pose detection, ResNet blocks for feature integration, and a 3D CNN discriminator, the model produces high-quality videos that maintain both spatial and temporal consistency. After 200 epochs, the model achieved an adversarial loss of 0.2265, with metrics like PSNR at 20.246, SSIM at 0.867, and LPIPS at 0.178. The high PSNR and SSIM values, along with the low LPIPS, show that the generated frames are well aligned and preserve important details. These results highlight the model's strong performance over time, consistently generating visually convincing videos of human activities using a reference image and source video. The model effectively transfers motion from video to image, creating realistic videos of human activity in comparison to existing models.
The convergence of deep learning and the Metaverse represents a pivotal frontier in the evolution of intelligent digital ecosystems. This paper presents a comprehensive survey of how deep learning techniques spanning convolutional, generative, transformer-based, and reinforcement architectures collectively enable perception, creation, cognition, and governance within immersive virtual worlds. Building upon this synthesis, we propose the Deep Learning-Empowered Metaverse Intelligence (DL-MI) framework, which unifies sensory intelligence, generative world-building, adaptive reasoning, and ethical-social governance into a cohesive architecture. The study illustrates how deep learning facilitates realistic avatar synthesis, dynamic environmental rendering, emotion-aware interaction, and predictive personalization, thereby transforming the Metaverse from reactive systems to anticipatory, self-evolving spaces. Key challenges such as data privacy, algorithmic bias, and computational sustainability are critically examined alongside emerging paradigms, including quantum-augmented AI and federated collaboration. By integrating technical, ethical, and societal dimensions, this survey provides a structured foundation for developing scalable, transparent, and human-centered Metaverse intelligence.
The mass-spring model is popular for representing the dynamics of distance-based systems. In particular, elastic materials can be easily simulated using the mass-spring model, thanks to the simplicity of its spring-length constraint. The model is commonly defined in a typical simulation domain, that is, 2D or 3D Euclidean space. If the elastic body is heavily deformed, however, its mass-spring configuration easily becomes hard or impossible to resolve in the typical domain. In order to tackle the challenge, we propose a dimension expansion method, which utilizes auxiliary coordinates for defining the model. By solving the problem in such an expanded domain, the solver can untangle complex spring configurations that are otherwise locked in local minima. Our study demonstrates the potential of the dimension expansion method for mass-spring-based elastic-body simulations and for different applications such as multi-dimensional data embedding problems.