
Omni-directional treadmills (ODTs) are pivotal for enabling natural locomotion in virtual environments, with applications in numerous fields. Despite the development of various ODT designs and applications over the decades, a comprehensive synthesis of ODTs’ technological evolution, applications, and challenges has not been previously established. In this article, we conduct a review of the advancements in the mechanisms, sensors, control strategies, auxiliary functions, and applications of ODTs through a thorough search of literature databases and commercial websites. Our analysis identifies a total of 32 distinct types of ODTs including the classic designs and the latest innovations. Among them, low-friction treadmills and XY-axis treadmills are the most prevalent ODTs. Our further analysis reveals that ODTs have unique advantages in many application scenarios, due to their abilities to enable infinite walking within a confined space with realistic full-body walking sensations. Additionally, we discuss several challenges and directions for the improvement in mechanism design, immersive experience, sensing technologies, and application scenarios. This review aims to provide researchers and engineers with a comprehensive overview of ODTs and to identify trends and challenges for guiding their future development in both academia and industry.
The increasing adoption of autonomous systems has highlighted the need to understand and predict users’ trust in these technologies. Existing trust prediction models perform well for most users but fail to capture the volatile, highly fluctuating trust patterns of “oscillators.” This raises two key challenges: First, developing a model that can accurately approximate oscillators’ dynamics, and second, identifying which individuals are likely to be oscillators so the model can be selectively applied. To address the first challenge, we propose a discounting trust model that prioritizes recent interactions while attenuating older ones, improving prediction accuracy for oscillators (optimal discount factor $\lambda = 0.8$). To address the second, we develop a classification model based on seven personal characteristics to screen for potential oscillators. Finally, we integrate the two approaches, applying the discounting trust model only to individuals predicted as oscillators. We used two datasets ($N = 130$ development and $N = 41$ validation). The development dataset was used to identify the optimal discount factor and train the screening classifier. The validation dataset was used to identify potential oscillators and to validate the performance of the discounting trust model on them. Four out of the 41 participants were identified as potential oscillators. We compared the discounting and nondiscounting baseline models using a linear mixed-effects model that accounted for autocorrelation. The discounting model significantly reduced prediction errors compared to the baseline model ($p < .001$). The findings showcase the value of incorporating both personal traits and a discount factor to enhance trust prediction for oscillators.
Rapid advances in artificial intelligence (AI) have reshaped traditional human–machine systems, giving rise to AI-powered Human–Machine Systems (AI-HMS). In AI-HMS, linear arbitration, the predominant shared control approach inherited from traditional human–machine systems, suffers from a performance-limiting linear structure and vulnerability to unbounded AI decision bias, potentially resulting in dangerous arbitration outcomes. To address these issues, we propose a Fused Distributional Nonlinear Arbitration (FDNA) approach, which fuses the probability distributions of human and machine strategies to construct a unified strategy distribution that effectively captures the nonlinear relationships between their decisions. The unified distribution is optimized with the conditional expected return as the objective, thereby enhancing arbitration safety while maintaining overall system performance. Through rigorous theoretical analysis, we demonstrate that FDNA offers advantages in terms of multipeak adaptability and robustness. Its effectiveness is further validated through comprehensive experiments in both discrete and continuous action spaces, under scenarios with and without intent inference and across different disturbance levels. The results demonstrate significant performance improvements over linear arbitration and other representative baseline methods.
Cardiopulmonary resuscitation (CPR) is a critical life-saving procedure, yet optimizing instructional design and accessibility in training remains an ongoing challenge. Current systems mostly rely on visual and audio cues, while haptic feedback remains largely unexplored in the literature. To address this gap, we developed MA-CPR, a multimodal assistance system that integrates visual, audio, and haptic feedback within a single CPR training platform. We systematically evaluated how CPR performance changes relative to a full multimodal feedback baseline when some specific feedback channels are removed, focusing on compression depth (via visual and/or haptic feedback), compression rate (via visual and/or audio feedback), and overall adherence to resuscitation guidelines. Regarding compression rate, participants received performance-based feedback through visual and/or audio modalities, and our results show that visual feedback remains the most reliable modality. Audio feedback proved effective for temporal guidance but less consistent for spatial parameters. Regarding compression depth, participants received performance-based feedback through visual and/or haptic modalities. Our results showed that although haptic feedback alone did not achieve the same compression depth as visual feedback, it still offers a useful alternative channel for conveying performance information when vision is restricted. Overall, participants did not perform significantly better when modalities were combined over a single modality, suggesting potential redundancy across channels. These findings indicate that haptic and audio feedback can meaningfully complement traditional training and even enable scenarios where visual feedback is not feasible, such as for visually impaired individuals. MA-CPR, thus, represents a step toward inclusive, multisensory CPR education that broadens accessibility while enhancing training flexibility.
Transfer learning has been widely applied in steady-state visual evoked potential (SSVEP) brain-computer interfaces (BCIs) to reduce calibration effort. However, most existing cross-subject transfer methods still require offline calibration data from the target subject for optimal SSVEP recognition. In this study, a novel cross-subject transfer method is proposed, which only a single-trial SSVEP signal from the target subject is required through an online adaptation mechanism. This method leverages shared temporal dynamics in early visual processing, including common impulse responses and transferred templates. Specifically, the common impulse response is derived by decomposing the source-subject SSVEP templates using linear superposition theory, while the transferred templates are constructed by weighted summation of the source-subject templates. These common impulse responses and transferred templates are then used to extract features from the target-subject SSVEP signals. Four correlation coefficients are computed to form the feature vector for SSVEP recognition. Experimental results on three widely used public SSVEP datasets show that the proposed method improves cross-subject transfer performance compared to state-of-the-art methods, demonstrating its ability to enable target subjects to use SSVEP-based BCIs without offline calibration.
Ensuring high-quality laparoscopic videos (LV) holds paramount significance during laparoscopic surgeries. In laparoscopy, the surgeons heavily rely on video feeds for visualization during laparoscopic procedures. These video feeds undergo different processing stages, and the loss of crucial information due to distortions and compression artifacts profoundly impacts the surgical efficacy and can lead to adverse patient outcomes. Therefore, laparoscopic video quality assessment (LVQA) has become crucial in monitoring the perceptual quality of LVs. This study aims to develop a “completely blind” no-reference (NR) LVQA model based on statistical properties of luminance, color, and structure maps. The quality discerning features of the aforementioned cues are computed using three modules. The first module computes the multifrequency information of luminance maps of the LV frames based on the entropy values of the Haar wavelet subbands. The second module extracts the chromatic deviations across the LV frames based on generalized Gaussian distribution model parameters. Finally, the structural variations among the LV frames are modeled based on the statistical parameters of the Weibull distribution. These features are computed on the individual frames of the LV and averaged across the number of frames to estimate the quality discerning feature vector of LV. Further, the likelihood scores of each feature map are computed with respect to pristine multivariate Gaussian model parameters, and the scores are pooled to estimate the final perceptual quality of LV. The proposed LVQA model does not require prior training, yet the model delivered superior performance compared to the off-the-shelf full reference and NR approaches.
Vascular interventional surgery serves as an effective treatment for vascular diseases. However, it demands exceptionally high-level skills from interventionists. Currently, novice interventionists face limited training resources and high time costs in skill acquisition, delaying their clinical readiness, while experienced interventionists endure substantial clinical workloads. To address these challenges, this article proposes a solution for the surgical skill learning and training of novice interventionists. First, in terms of system development, we designed a training system for novice interventionists’ surgical skills. It consists of an interactive device adapted for clinical operations and a high-fidelity virtual reality surgical training scenario. Second, in terms of surgical skills learning and training, we studied the influence of different design factors on the surgical skills learning and training of novice interventionists. The proposed system’s intuitive structural design and high-fidelity haptic feedback not only facilitate surgical skills learning but also enhance the operating experience, supporting a better understanding of surgical procedures. Our solution holds great potential in assisting novice interventionists in their surgical skills development.
This article addresses the vulnerability of player strategies in continuous action-iterated dilemma (CAID) games to external interception. To mitigate this risk, a privacy-preserving mechanism based on an output masking function is proposed within the CAID framework. This mechanism safeguards initial player strategies, preventing foreign entities from accurately determining them and manipulating the game outcome. The protection of players' initial strategies ensures the integrity of the average consensus value. In contrast to existing methods that add random noise to players' strategies, the proposed approach uses an output masking technique to conceal players' states, making them indiscernible from others. This ensures that all players in the CAID game converge exactly to the average of their initial strategies while effectively safeguarding the confidentiality of their initial strategies. The proposed methodology is supported by a detailed theoretical analysis of the consensus process. Moreover, the effectiveness of the proposed strategy is demonstrated through simulations of two evolutionary game examples and comparative performance analysis against existing privacy-preserving methods.
Human activity recognition based on radar faces several challenges. First, the strong scattering echoes from the torso often conceal the weaker scattering echoes from the limbs, head, and other body parts, thereby limiting the expression quality of human activity regularity. Besides, the lack of a large number of labeled radar data samples limits the performance of the network model when training. To overcome these issues, a semi-supervised human activity recognition method based on multichannel scattering separation is proposed. First, principal component analysis is leveraged to separate weakly scattered limb signals based on multichannel radar signals while filtering out strongly scattered torso signals, to highlight detailed information about human activity. In addition, a semi-supervised learning approach is explored to effectively utilize unlabeled data, enhancing model precision without extensive labeled samples. During semi-supervised learning, two models in parallel are trained. Subsequently, by leveraging the differences between the two models and making predictions under different perturbations, pseudolabels are generated when the prediction results are consistent and the confidence is greater than the threshold, stored in the queue structure and added to the future training in the next parameter update. Experimental results demonstrate that the proposed method consistently outperforms supervised baselines and state-of-the-art semi-supervised methods, particularly under severe label scarcity. With only six labeled samples, it achieves 61.59% accuracy for six human activities, surpassing MixMatch and SoftMatch by 40.94% and 11.03%, respectively.
With sleep health issues, such as sleep disorders, gaining increasing attention, precise assessment of sleep quality is crucial for improving work efficiency and quality of life. Existing research utilizes the spatiotemporal features of multichannel brain signals and the spatial topological information between brain regions, but the fundamental problem of low accuracy in sleep stage classification persists. To address this, this work proposes a multiview signal integration network based on manifold learning for sleep stage classification. First, three types of feature representations for multiphysiological signals are constructed through short-time Fourier transform, wavelet transform, and raw signal processing, respectively. Second, a simple attention module fusing Euclidean space with manifold geometry is designed to represent and enhance these multiphysiological modalities in Riemannian manifold space, mining potential spatiotemporal representations. Building upon this, the temporal dependencies of these three types of sleep features across adjacent time periods are modeled using long short-term memory to accomplish sleep stage classification. Extensive experimental results on the public ISRUC-S1 and ISRUC-S3 sleep datasets confirm that the feature combination approach and manifold attention can effectively improve the classification accuracy.
This study investigates how matched and mismatched visual and auditory stimuli in virtual reality (VR) influence stress perception and immersion. Utilizing VR's capacity for controlled multisensory experiences, we explored whether stress-related physiological responses can serve as an indicator of immersion quality. Twenty-two participants were exposed to both forest and city VR environments with either congruent (matched) or incongruent (mismatched) audiovisual stimuli. Using a 2 × 2 within-subject design, each participant thus experienced four conditions (city_ref, city_conf, forest_ref, forest_conf) with passive navigation. Physiological stress responses were measured using electroencephalography (EEG) and blood volume pulse. Subjective stress levels were assessed using the State-Trait Anxiety Inventory and NASA Task Load Index frustration subscale, whereas immersion was assessed using a validated presence questionnaire. Results indicated that matched sensory inputs in the forest environment reduced stress and enhanced immersion, while matched inputs in the city environment increased stress but also enhanced immersion. In mismatched conditions, immersion was reduced due to the sensory conflict. Calming sounds mitigated stress even with stressful visuals, whereas disruptive sounds increased stress despite calming visuals. This dominance aligned with existing findings on the dominant role of auditory input in stress perception. EEG features, particularly frontal asymmetry in the alpha and gamma bands, were sensitive to both stress and immersion. The novelty of this work lies in linking sensory congruence, stress-related physiological responses, and immersion. The findings highlight the importance of proper multisensory integration in VR design. These insights have practical implications for the design of therapeutic VR applications.
Motor dysfunction caused by stroke, Parkinson's disease, and traumatic injuries seriously affects patients' quality of life. Traditional rehabilitation methods often suffer from limited accessibility, high costs, and inefficient data management, especially in remote rehabilitation scenarios. To overcome these issues, we present a cross-platform remote rehabilitation system that integrates virtual reality (VR) for patient-side training, augmented reality (AR) for clinician-side observation, dynamic digital twin (DT) for motion representation, smart treadmill feedback for physical interaction, and blockchain-based mechanisms for trusted data management. The proposed system achieves a closed-loop rehabilitation framework that enables low-latency multimodal synchronization, DT-driven feedback, and interactive remote intervention between patients and clinicians. The experimental results show that the average local perception latency from skeleton tracking to recognition output is approximately 230 ms; the average end-to-end clinician-view latency, covering skeleton recognition, network transmission, DT mapping, and remote rendering, remains below 400 ms; the success rate of clinician-side intervention operations is 96%; the on-chain write success rate is 100%; and the average authorization latency is 89 ms. This work provides a secure and responsive framework for remote rehabilitation by combining immersive VR-enabled training, AR-assisted clinical interaction, DT-based motion modeling, and trustworthy data management, offering practical support for personalized and privacy-preserving rehabilitation services.
Rapid manufacturing evolution accelerates product iteration, requiring collaborative robots to rapidly adapt and deploy new action programs. Although large language models (LLMs) can generate logically consistent instructions, they lack contextual understanding of human-robot collaborative environments. We propose an understanding and perceiving the environment with requests model that fuses multimodal information to enable embodied robots to interpret user intent within environmental context. In addition, a flexible task planning framework for multiagent systems is introduced, allowing agents to form flexible teams according to task requirements. This framework grants agents reflective capabilities, enabling them to analyze failures and replan accordingly, thus enhancing task planning flexibility. Appropriately designed prompts for embodied intelligence robots are used to mitigate the hallucination effect of LLMs, further improving the accuracy of task planning when using LLMs. In collaborative aerospace component assembly experiments, the proposed system significantly improved natural-language-driven task execution, achieving a 10.80% improvement in final success rate.
Human cognitive performance is shaped not only by vision, hearing, and workload but also by latent neuroendocrine stress responses. Conventional unimodal physiological sensing, i.e., relying on a single biosignal modality such as ECG alone or breath-VOC sensing alone, captures only partial stress physiology and is vulnerable to noise and confounds, limiting real-time assessment in human-machine systems. To address this gap, we present an olfactory-enhanced cardiac sensing platform that integrates a multichannel MOS/MEMS e-nose with a compact ECG front-end, enabling synchronized acquisition of metabolic and cardiovascular signals. On top of this hardware, we propose 1D_Att_CNN, a dual-modal fusion model with parallel 1-D-CNN encoders and squeeze-and-excitation network (SENet) attention, which dynamically weights fast cardiac rhythms and slower VOC kinetics. Evaluated on 1215 trials from 27 participants, the system achieves 93.24% accuracy, surpassing SVM, XGBoost, and ensemble baselines. Hardware validation confirms stable ECG capture and stress-sensitive VOC responses. This compact hardware-algorithm pipeline establishes a robust, noninvasive solution for monitoring psychological stress, advancing human-factor engineering, cognitive-state assessment, and situational awareness in complex environments.
In recent years, single-image-based 3-D human modeling has become a promising research direction, with potential applications in augmented reality, human-computer interaction. However, 3-D human body modeling from a single image often produces physically unreasonable results in close-interaction scenarios due to severe occlusion and depth ambiguity. We propose a physics-aware framework for holistic human mesh recovery (Phy-HHMR), a physics-driven holistic reconstruction system that recovers expressive 3-D meshes within a unified framework. Our method features an interaction-aware optimization pipeline, combining physical consistency priors with spatially guided diffusion refinement. By employing geometric gap regularization and tokenized motion memory, the system ensures nonpenetration and contact plausibility, maintaining kinematic accuracy and a physical foundation. Phy-HHMR leverages a system-level joint optimization strategy and paired-flow couplers to sustain overall mechanical consistency. Simulation results show that our method achieves state-of-the-art performance and stronger robustness to reconstruction defects, providing a reliable basis for virtual reality and human-computer collaboration.
The growing popularity of virtual reality (VR) provides new opportunities for cultural heritage protection and online education. Many studies have investigated how to guide users to move to a specific view/position. However, the coguided interaction of position and gaze is still under researched. In this article, a novel interaction model named GazeTance Guidance (GTG) is proposed. This model utilizes head rotation and viewing distance to improve the effect of room-scale user guidance. To explore the efficiency of GTG, we conducted two within-subject studies with 25 and 31 participants navigating through virtual museums: first, we evaluated user acceptance of audio commentary and virtual guidance markers in terms of presence and user experience, and assessed the feasibility of memory tasks. second, we separately investigated the effects of gaze and distance guidance on motion sickness, cognitive load, memory efficiency, and user behavior. The experimental results prove that the guidance markers do not affect user experience and sense of presence. Compared with independent effects, the combination of gaze and distance guidance can reduce the user's motion sickness and significantly improve users' memory effect. These findings provide new inspiration for the interaction design of complex VR tours.
The automation of reporting in the manufacturing sector, enabled by advanced sensing technologies and sophisticated software, offers considerable promise for enhancing operational efficiency and profitability. Despite substantial economic benefits, fully realizing automated reporting remains hindered by critical challenges in human action recognition. These obstacles include visual occlusions in dynamic environments, heavy reliance on large-scale labeled video datasets, requirements for real-time activity identification, and the complexity of recognizing sequential and interdependent tasks typical in manufacturing settings. This article introduces a novel CNN-based framework specifically designed to overcome these challenges, thereby enhancing automated reporting capabilities in manufacturing. Our proposed method incorporates innovative geometric convolution layers, significantly improving the accuracy of motion identification and posture estimation under challenging conditions. To reduce dependence on extensive labeled datasets, we integrate a self-supervised learning strategy, leveraging unlabeled video data effectively. Moreover, we propose a robust dual-channel graph convolutional network that concurrently addresses multitask action recognition and prediction, effectively capturing complex sequential activities. Collectively, these contributions equip manufacturers with the tools necessary for precise, real-time monitoring of worker actions, enabling optimized operations, enhanced worker safety, and ultimately, improved profitability.
A human-machine interface (HMI) is an essential component of any modern industrial machine, acting as the portal for human workers to setup, monitor, control, and troubleshoot machines. Yet, existing HMIs often pose challenges in usability and learnability due to overwhelming data, nonstandard interfaces, and physical design limitations. We argue that these challenges stem from a lack of spatiotemporal alignment of data and controls with the machine, leading to information overload and loss as the operator repeatedly translates signals between the HMI and machine. We propose a novel augmented reality (AR) interface integrated with Internet of Things (IoT), which affords spatially aware, natural mapping of real-time data and controls with the machine while providing visibility and immediate feedback. We build the IoT-integrated AR system through a Kepware-ThingWorx-Vuforia pipeline implemented in a model cyber-physical factory for cellphone assembly. We conducted a between-subjects user study with 20 participants split into two groups, who operated three assembly line stations using our system deployed on HoloLens 2 and the machine's built-in Siemens HMIs in opposite orders. The findings show the significant effects of our AR interfaces on improving efficiency, lowering task completion times, reducing errors, and enhancing usability and mental load. We conclude the paper by discussing the limitations and several directions for future research on IoT-integrate AR interfaces to transform how humans interact with industrial machines.
Human-in-the-loop switched systems often suffer from unmeasurable states and strict safety constraints, which may lead to unacceptable workload or even catastrophic errors for human operators. This article proposes an adaptive-enhanced constraint-management strategy for such human-machine shared-control platforms. The main results and contributions are as follows. first, a self-adjusting event-triggered mechanism is introduced to reduce the update frequency of the autopilot, alleviating the operator's cognitive burden while ensuring communication efficiency; Second, an adaptive state observer is designed to reconstruct unmeasurable states in real time, providing accurate state information for feedback control; Third, by incorporating a tan-type barrier Lyapunov function, we theoretically guarantee that all state variables strictly remain within predefined human-safe ranges throughout system operation; fourth, fuzzy-logic systems are employed to approximate unknown nonlinearities arising from pilot-autopilot interactions, with established approximation error bounds; Finally, under the average dwell-time switching law, we rigorously prove that the closed-loop human-machine system achieves uniformly ultimately bounded stability while excluding Zeno behavior. Numerical simulation results validate the effectiveness of the proposed approach, demonstrating improved tracking accuracy and reduced control update frequency compared with traditional periodic control schemes.