Traditional Chinese Medicine (TCM) massage is well known for its therapeutic effectiveness; however, its training and dissemination are often hindered by a lengthy learning process and high skill-entry requirements. To alleviate these challenges, this paper presents the design and initial implementation of an Augmented Reality (AR)–based system for learning and practicing TCM massage. The proposed system integrates real-time automatic acupoint recognition and calibration, along with visualized feedback of hand-applied pressure on Quest devices, to support more intuitive and guided practice. A preliminary user study was conducted to evaluate the feasibility of the system, through which its current limitations and potential directions for improvement were also identified. The results indicate that incorporating AR technology into TCM massage education can enhance learning efficiency and learner confidence, while providing effective assistance during practical operations. This work demonstrates the potential of immersive technologies to support the modernization and dissemination of TCM massage training.
In high-risk and high-complexity supervisory control tasks, the design of system transparency involves an antagonistic trade-off between trust gain and cognitive load. To elucidate the mechanism by which transparency styles influence collaboration efficiency, this study proposes the Information Effectiveness–Cognitive Hindrance Dual-Path Competition Model (TCDP). This model aims to reveal the dynamic interplay between trust gain and cognitive hindrance in transparency design, while systematically examining the serial mediation effect of information interaction style combinations on collaboration efficiency. By constructing an information entropy quantification framework and a dynamic threshold model, a system test involving multi-modal and multi-strategy combinations was conducted with 84 participants across three experimental tasks (equipment monitoring, fault management, and multi-task collaboration). The results indicate that: (1) Optimal information combinations significantly enhance trust by improving transparency (disclosure, clarity, accuracy), thereby boosting collaboration efficiency; (2) Conversely, certain information combinations exacerbate cognitive load by increasing information entropy (modal entropy, interaction entropy), thus inhibiting efficiency; (3) Task complexity moderates the weights of these dual paths, with the hindrance path exerting a stronger effect in high-complexity tasks; and (4) Collaboration efficiency follows an inverted U-shaped curve relative to the Transparency-to-Entropy ratio, indicating the existence of a task-dependent optimal threshold. Theoretically and empirically, this study unveils the dual-path "trust–load" competition mechanism, providing quantitative grounds and design principles for building adaptive human-machine systems characterized by high trust and low cognitive load.
The information symmetry between humans and machines can enhance mutual perception and understanding, leading to more robust cooperation. Intention recognition is a key technology in natural human-machine collaboration (HMC). However, as the complexity of the system increases, the amount of information and the types of tasks become numerous, which leads to continuous dynamic changes in interaction intentions. New sensing technologies such as electroencephalogram (EEG) have provided a continuous and unobtrusive monitoring approach to accurately and effectively identify intentions. But the complexity of physiological responses and the uncertain nature of intention make cross-subject recognition difficult, resulting in poor generalization performance and coarse-grained recognition patterns. To address these limitations, we proposed a framework for modeling tasks in complex systems and applied it to model tasks in an industrial system. Then, we use operators' EEG data to effectively recognize the fine-grained intention patterns within different typical task scenarios, such as monitoring production and communication tasks. By inputting the improved multi-channel phase synchronization features into a machine learning classifier, cross-subject accuracy rates of 99.42% and 99.91% were achieved. This work furnishes systematic, field-tested cases for task modeling in the industrial field, demonstrates high-performance implicit intention recognition with a single EEG modality, and refines the granularity of implicit intention recognition. It provides theoretical underpinnings and technical support for both human-machine information symmetry and the advancement of HMC hybrid intelligence.
Accurate object scaling is essential in building 3D scenes, yet manual adjustment is time-consuming for developers and hard to calculate, and 3D models imported from various sources often lack consistent scaling. To address this challenge, we propose an image-similarity-based scale estimation method that can automatically predict appropriate dimensions for 3D models using geometric and visual embeddings. The method shows superior accuracy and speed compared to baseline depth-based approaches. It offers an efficient solution for automated scene generation and model normalization.
Large Language Models (LLMs) are being used more and more extensively for automated evaluation in various scenarios. Previous studies have attempted to fine-tune open-source LLMs to replicate the evaluation explanations and judgments of powerful proprietary models, such as GPT-4. However, these methods are largely limited to text-based analyses under predefined general criteria, resulting in reduced adaptability for unseen instructions and demonstrating instability in evaluating adherence to quantitative and structural constraints. To address these limitations, we propose a novel evaluation framework, ARJudge, that adaptively formulates evaluation criteria and synthesizes both text-based and code-driven analyses to evaluate LLM responses. ARJudge consists of two components: a fine-tuned Analyzer that generates multi-faceted evaluation analyses and a tuning-free Refiner that combines and refines all analyses to make the final judgment. We construct a Composite Analysis Corpus that integrates tasks for evaluation criteria generation alongside text-based and code-driven analysis generation to train the Analyzer. Our results demonstrate that ARJudge outperforms existing fine-tuned evaluators in effectiveness and robustness. Furthermore, it demonstrates the importance of multi-faceted evaluation and code-driven analyses in enhancing evaluation capabilities.
Holography allows realistic 3D visualizations, offering lifelike representations visible to the naked eye. This study presents a multimodal interface framework integrated into a holographic map (HM) system developed for large shopping malls. A total of 56 college students and staff voluntarily participated in interaction tasks. Experiment 1 involved 20 randomly selected participants (10 males and 10 females, aged 20-30) who navigated the HM using seven predefined behavioral instructions. Experiment 2 included 40 participants (aged 20-35) with prior multimodal interaction experience. It evaluated speech, gesture, and eye-tracking modalities across the same instructions. To maintain consistency, "eye-tracking" was used throughout to describe gaze-based input. Participants assessed interaction quality using criteria such as smoothness, fatigue, acceptance level, matching degree, and operability. The findings provide insights into selecting effective interaction modes for HM systems, enhancing user experience and interaction efficiency in public spaces.
This study focuses on the innovative application of HCI and XR technologies in behavioral skills training (BST) in the digital age, exploring their potential in education, especially experimental training. Despite the opportunities these technologies offer for immersive BST, traditional methods remain mainstream, with XR devices like HMDs causing user discomfort and current research lacking in evaluating user experience. To address these issues, we propose the spatial reality display (SRD) method, a new BST approach based on spatial reality display. This method uses autostereoscopic technology to avoid HMD discomfort, employs intuitive gesture interactions to reduce learning costs, and integrates BST content into serious games (SGs) to enhance user acceptance. Using the aluminothermic reaction in chemistry experiments as an example, we developed a Unity3D-based XR application allowing users to conduct experiments in a 3D virtual environment. Our study compared the SRD method with traditional BST through simulation, questionnaires, and interviews, revealing significant advantages of SRD in enhancing user skills and intrinsic motivation.
Mental health is a critical global issue, with many individuals lacking timely psychological support due to limited resources and high costs. Traditional interpretations of the House-Tree-Person (HTP) test, which rely on subjective judgment, are often time-consuming and lack scalability, necessitating more objective and efficient methods. This study explores the application of artificial intelligence (AI) technology, especially large language models (LLMs), to automate HTP test analysis and provide personalized therapeutic interventions, aiming to enhance the efficiency and accuracy of psychological assessments. We constructed a diverse HTP drawing database and developed an AI-powered system that integrates HTP analysis with music, visual, and aroma therapies. A user experiment with eight participants was conducted to evaluate the system's performance through a Likert scale questionnaire and semi-structured interviews. Results showed that participants had high trust in the AI-assisted HTP test results, with an average satisfaction score of 3.75 for the overall effectiveness of art therapy. This demonstrates the potential of AI to enhance both assessment accuracy and therapeutic outcomes. Future work will focus on improving model interpretability, exploring ethical implications, and expanding the application scope to other psychological tools to achieve more effective and sustainable treatment outcomes.
With the rapid advancement of Virtual Reality (VR) technology, the gaming industry is confronted with unprecedented opportunities and challenges. Immersion, as a fundamental aspect of the VR gaming experience, has a direct impact on player engagement and satisfaction. Consequently, the design of game elements that can enhance immersion has become a critical factor in improving the overall quality of games. This study, based on the Kano model and the Entropy Weight TOPSIS method, examines the relationship between immersion elements and user satisfaction in VR game design. The Kano model categorizes game design elements into Must-be Attributes, One-dimensional Attributes, Attractive Attributes, Indifferent Attributes, Reverse Attributes, and Questionable Attributes, thereby enabling designers to identify and classify different player needs and optimize the design of immersive experiences. Simultaneously, the Entropy Weight TOPSIS method combines the entropy weight approach with the TOPSIS method to objectively calculate the importance and weights of various design elements, evaluating the relative effectiveness of different design alternatives, and ultimately determining the immersion design scheme that best aligns with player preferences. The findings indicate that immersion and user satisfaction do not always exhibit a positive correlation, as factors such as interaction usability, ease of use, comfort, gameplay content, game difficulty, graphics and sound effects, stability, and social interactivity also contribute to overall satisfaction. Based on these insights, this paper constructs a comprehensive immersion element library for VR game mechanics, and applies the Kano model and Entropy Weight TOPSIS method to evaluate and prioritize the design elements. The selected elements are those that most effectively fulfill both immersion and user satisfaction requirements, providing a scientific foundation for VR game design.
In supervisory control tasks, particularly in high-risk fields, operators need to collaborate with automated intelligent agents to manage dynamic, time-sensitive, and uncertain information. Effective human–agent collaboration relies on transparent interface communication to align with the operator’s cognition and enhance trust. This paper proposes a human-centered adaptive transparency information design framework (ATDF), which dynamically adjusts the display of transparency information based on the operator’s needs and the task type. This ensures that information is accurately conveyed at critical moments, thereby enhancing trust, task performance, and interface usability. Additionally, the paper introduces a novel user research method, Heu–Kano, to explore the prioritization of transparency needs and presents a model based on eye-tracking and machine learning to identify different types of human–agent interactions. This research provides new insights into human-centered explainability in supervisory control tasks.
Holographic displays represent a significant application of future human-computer interaction, with desktop-based digital light field holography already seeing widespread use. As this technology develops, exploring 3D interaction design strategies and user experience research specific to holographic displays becomes essential. Despite the rapid advancements in display technology, there are still significant gaps in compatible interaction paradigms. The mortise and tenon structure, an intangible cultural heritage of China, combines concepts from philosophy, material science, mechanics, mathematics, and aesthetics. It has unique characteristics of traditional culture and spatial organization, making it a worthwhile object for developing an experience system that combines craftsmanship, art, and culture through holographic 3D display and interaction. This paper reviews the historical development, research methodologies, and frameworks for human-computer interaction interfaces. It focuses on designing interaction interfaces within holographic environments, incorporating the Chinese cultural heritage of the mortise and tenon structures. The study constructs an interactive experience system using the Unity platform, Leap Motion for gesture recognition, and a Sony holographic display screen for visualization. The project evaluated the overall performance of this holographic display interaction platform through comparative experiments (N = 16). The results demonstrate that the platform offers a positive play experience for structural cognition. The project provides valuable user data and design process as a case study reference, contributing to the education field of the holographic interaction.
This paper introduces the overall framework of the intelligent electrical sporadic materials management system based on the ubiquitous Internet of Things. The system includes storage equipment and an intelligent management terminal. It makes optimal design for storage equipment and intelligent management terminals. The dynamic modeling of a three-layer supply chain system of power materials is constructed and composed of manufacturers, suppliers, and raw materials. The inventory control strategy of each link in the supply chain of power materials under different cooperation levels is studied by introducing the corresponding adjustment variables. Through the experiment and analysis of the system, it is proved that the system has good performance in label management, resource management, inventory management, disposal management, system management and essential management. The number of concurrent users, response speed, stability and other aspects of the system are excellent. The database is complete, independent, and secure. And the data is objectively reasonable and repairable. The system can lay a specific technical foundation for intelligent management of electrical sporadic materials.
The industrial field has entered the era of industrial intelligence,following the developmental stages of mechanization and informatization.The hallmark of the new industrial revolution is cyber-physical systems(CPS),which merge the digital and physical worlds.However,owing to advancements in various fields,a fully autonomous system is not achievable in the near future.In the process of deepening industrial intelligence,leveraging their respective advantages to achieve safe and efficient cooperation remains a significant issue.Many studies have shown that accidents caused by human error constitute a significant portion of safety-centric complex systems,with mental workload being the primary factor leading to such errors.As human-in-the-loop research based on transparent interfaces,such as the brain-computer interface,advances,passively and naturally integrating human cognitive models into human-machine systems has become a trend,thereby participating in the decision-making and control of the future.Mental workload is the key cognitive component related to human cognition in safety-centered complex systems research.As an important factor reflecting system performance,it is crucial to evaluate mental workload scientifically and quantitatively in future industrial human-cyber-physical systems research.Utilizing this as the starting point,this paper comprehensively reviews the development of the theoretical basis of mental workload,research progress of theoretical framework,development of mental load assessment,existing experimental paradigms,and induction methods.By visiting related factories and communicating with automation experts and job operators,the components and characteristics of industrial CPSs are analyzed,and the characteristics and core operation mode of the human-computer interaction(HCI)system of industrial CPSs are proposed.Three primary task types in the industrial context have been identified,and an experimental paradigm for mental workload research,tailored to the industrial environment,has been established.Thus far,a simulated task load paradigm has been applied in the field of industrial systems.In the modeling phase,the key electroencephalogram(EEG)sensitivity parameters for three tasks must first be extracted.Among the features that showed significant performance under different mental workload levels,the characteristics of absolute EEG power exhibit the best performance in distinguishing mental workload.Subsequently,multitype task recognition models based on EEG signals are established.Results show that in monitoring and verification,control operation,and communication scenarios,the test accuracy rates of the algorithm model were 88.14%,94.72%,and 82.42%,respectively.Through five-fold cross-validation and mesh parameter optimization,the optimal parameters for the model are obtained.To assess the model's generalization ability,a subject-wise method is employed.Upon examination,the average recognition accuracy of the random forest model based on absolute EEG power features reached 97.85%,96.95%,and 89.88%,and the highest recognition accuracy can reach 100%.The nonlinear entropy features in the frequency domain showed crossover,while the fusion features did not show an obvious trend beyond the absolute EEG power features.This study provides a theoretical basis for the elucidation of the relationship between operator physiology and system workload,improving the productivity of industrial CPS and operator enthusiasm and providing a new perspective for promoting the research on human-in-the-loop based on transparent interfaces,natural human-computer interactions,adaptive software,and brain-computer interfaces.
Acupoint detection plays an important role in intelligent acupuncture, rehabilitation training, and the digitalization of traditional Chinese medicine. However, due to the weak textures, dense distribution, and complex structural characteristics of back acupoints, existing localization methods largely rely on heuristic rules or manual annotation, leading to limited efficiency, robustness, and cross-subject generalization, and thus failing to meet the demands of large-scale applications in complex scenarios. With the rapid development of end-to-end detection frameworks and vision Transformers, sparse query-based point set modeling has provided new opportunities for building accurate and generalizable acupoint detection models. In this work, we propose AcupointDETR, an end-to-end Transformer-based framework for multi-acupoint localization. Acupoint recognition is formulated as a sparse query-driven keypoint detection task, enabling unified localization without post-processing. Considering the fixed number and well-defined spatial structure of acupoints, we design a Keypoint-Oriented Query Selection (KOQS) mechanism based on encoder response peaks to construct structured queries, and introduce an Acupoint-Oriented Denoising (AOD) strategy to enhance robustness against noise, weak textures, and local occlusions. Furthermore, a Dynamic-Acupoint Prior Gated Fusion (D-APGF) module is incorporated into the decoder to integrate acupoint structural priors, improving global layout consistency and localization accuracy. Experimental results demonstrate that AcupointDETR significantly outperforms existing methods on back acupoint recognition, and ablation studies further verify the effectiveness of the proposed modules in accelerating convergence and improving recognition accuracy.
Traditional computer-aided design (CAD) tools based on keyboard and mouse interactions present challenges to efficiency and quality in terms of free creation, iteration efficiency, and operational experience. This paper proposes a multimodal interactive guiding grammar framework for virtual-reality modeling scenarios. In studies on the characteristics of multimodal combinations, the present heuristic method leads to a high cognitive load and experimental cost, and the heuristic is difficult. Therefore, we propose a two-stage heuristic method. Subsequently, we use this method to decompose the interaction task into two stages: modalities and interactive operations. A clearer and more specific reference for multimodal selection is provided for the interaction design phase, based on the two dimensions of modality and task. This can reduce the heuristic cost and difficulty of users in completing tasks in multimodal interaction scenarios. Finally, the proposed virtual modeling platform which is used for experimental verification, and the results shows that, compared with traditional modeling, the interactive modeling method built according to the results of modal interaction in this study can better satisfy users in terms of emotional experience. In the index evaluation, high scores are achieved, with an average score of 5 points (1 and 7 are the lowest and highest scores of the evaluation index, respectively), and the large-area distribution of high scores also shows a good user experience. To a certain extent, this method improves the interaction mode of the previous 3D modeling.
In recent years, outdoor hiking has gained popularity, driven by increasing health awareness. However, enthusiasts of hiking encounter challenges related to equipment carrying, long-distance journeys, and environmental uncertainties. These challenges include the limitation of frequently checking routes on mobile devices due to the use of trekking poles. To address these obstacles, we undertake research and design efforts that integrate outdoor scenarios with Augmented Reality (AR) technology. Our approach involves establishing AR user interface design guidelines and design strategies tailored for outdoor scenarios, derived from a comprehensive systematic literature review. We propose a multimodal AR system solution that leverages AR glasses and trekking poles as dual inputs. The system incorporates multimodal interaction and natural information presentation, aiming to create a multilevel interactive experience enriched with tangible physical evidence. The goal is to establish a seamless connection between the virtual and real worlds, mitigating the challenges faced by hiking enthusiasts. Evaluation experiments conducted demonstrate that the system successfully meets user needs and expectations in terms of multimodal input, visual information recognizability, and the rationality of user interface layout. The system is noted for its close alignment with the usage scene, exhibiting enhanced user-friendliness in practical applications. This study provides valuable insights into the design of AR systems in outdoor sports scenes, offering a reference point for future endeavors in this domain.
This article explores the application of augmented reality technology in the field of traditional Chinese massage (TCM), aiming to provide users with an immersive and personalized massage experience at home. Through qualitative and quantitative research methods, this study identifies users’ preferences and pain points and proposes a series of innovative AR massage design solutions, including AR massage route guidance, personalized experience, and disease diagnosis. This study also analyzes the potential of AR technology in medical education and practice. The application of AR technology provides new possibilities for TCM, enabling users to enjoy a more realistic, comfortable, and personalized massage experience at home. Meanwhile, AR technology can also provide users with more accurate disease diagnosis and health management, help users understand their own physical condition, and prevent the occurrence and development of diseases through massage health care. In addition, we compared the influence of video guidance, static guidance, and dynamic guidance in the AR massage learning environment through experiments and emphasized the excellent performance of dynamic guidance in massage gesture restoration and the sensitivity of virtual and real combinations. The results of this study provide new ideas and methods for the learning and practice of TCM, and also provide useful references for the practical application of AR technology in the medical field.
Comprehensively understanding surgical scenes in Surgical Visual Question Answering (Surgical VQA) requires reasoning over multiple objects. Previous approaches address this task using cross-modal fusion strategies to enhance reasoning ability. However, these methods often struggle with limited scene understanding and question comprehension, and some rely on external resources (e.g., pre-extracted object features), which can introduce errors and generalize poorly across diverse surgical environments. To address these challenges, we propose SCAN, a simple yet effective memory-augmented framework that leverages Multimodal LLMs to improve surgical context comprehension via Self-Contained Inquiry. SCAN operates autonomously, generating two types of memory for context augmentation: Direct Memory (DM), which provides multiple candidates (or hints) to the final answer, and Indirect Memory (IM), which consists of self-contained question-hint pairs to capture broader scene context. DM directly assists in answering the question, while IM enhances understanding of the surgical scene beyond the immediate query. Reasoning over these object-aware memories enables the model to accurately interpret images and respond to questions. Extensive experiments on three publicly available Surgical VQA datasets demonstrate that SCAN achieves state-of-the-art performance, offering improved accuracy and robustness across various surgical scenarios.
Diffusion models have demonstrated substantial success in controllable generation for continuous modalities, positioning them as highly suitable for tasks such as human motion generation. However, existing approaches are typically limited to single-task applications, such as text-to-motion generation, and often lack versatility and editing capabilities. To overcome these limitations, we propose UniMotion-DM, a unified framework for both text-motion generation and editing based on diffusion models. UniMotion-DM integrates three core components: 1) a Contrastive Text-Motion Variational Autoencoder (CTMV), which aligns text and motion in a shared latent space using contrastive learning; 2) a controllable diffusion model tailored to the CTMV representation for generating and editing multimodal content; and 3) a Multimodal Conditional Representation and Editing (MCRE) module that leverages CLIP embeddings to enable precise and flexible control across various tasks. The ability of UniMotion-DM to seamlessly handle text-to-motion generation, motion captioning, motion completion, and multimodal editing results in significant improvements in both quantitative and qualitative evaluations. Beyond conventional domains such as gaming and virtual reality, we emphasize UniMotion-DM’s potential in underexplored fields such as healthcare and creative industries. For example, UniMotion-DM could be used to generate personalized physical therapy routines or assist designers in rapidly prototyping motion-based narratives. By addressing these emerging applications, UniMotion-DM paves the way for utilizing multimodal generative models in interdisciplinary and socially impactful areas.
Tracking the articulated poses of multiple individuals in complex videos is a highly challenging task due to a variety of factors that compromise the accuracy of estimation and tracking. Existing frameworks often rely on intricate propagation strategies and extensive exchange of flow data between video frames. In this context, we propose a spatiotemporal sampling framework that addresses the degradation of frames at the feature level, offering a simple yet effective network block. Our spatiotemporal sampling mechanism empowers the framework to extract meaningful features from neighboring video frames, thereby optimizing the accuracy of pose detection in the current frame. This approach results in significant improvements in running latency. When evaluated on the COCO dataset and the mixed dataset, our approach outperforms other methods in terms of average precision (AP), recall rate (AR), and acceleration ratio. Specifically, we achieve a 3.7% increase in AP, a 1.77% increase in AR, and a speedup of 1.51 times compared to mainstream state-of-the-art (SOTA) methods. Furthermore, when evaluated on the PoseTrack2018 dataset, our approach demonstrates superior accuracy in multi-object tracking, as measured by the multi-object tracking accuracy (MOTA) metric. Our method achieves an impressive 11.7% increase in MOTA compared to the prevailing SOTA methods.