
As the core carrier for achieving human-machine collaborative work in intelligent manufacturing scenarios, one of the key indicators of the bionic intelligence level of humanoid robots is the real-time and accurate perception of human emotional states driven by visual perception. Facial expression recognition (FER) is the core supporting element of this perceptual ability. The existing FER algorithm generally suffers from insufficient capture of facial region correlation, which makes it difficult to meet the practical needs of natural collaboration between humanoid robots and human operators. This is also one of the core scientific issues in the development of industrial humanoid robots that combine emotions and intelligence. This paper addresses this issue by combining a self-attention mechanism with a VGG16-BN network, proposing a novel visual perception network called SAVNet, and integrating it into the expression recognition system of humanoid robots. The proposed network embedding visual self-attention mechanism (VSAM) enables the model to accurately capture the associated features of key facial regions within the sensory field, effectively improving the FER performance of industrial humanoid robots in complex industrial environments. The evaluation results on the FER2013, RAF-DB, CK+, and AffectNet standard datasets show that the VSAM module has significant effects in optimizing the internal representation of features and enhancing the model's robustness against industrial scene interference. For example, after 50 rounds of training on the RAF-DB dataset, the highest classification accuracy of VGG16 on the RAF-DB test set was 80.07%, while the proposed model achieved an expression recognition rate of 82.26%, which was 2.19% higher than that of VGG16. The practical verification of applying SAVNet to the prototype system of humanoid robots shows that this method not only achieves leading recognition accuracy on standard datasets, but also can stably respond to human expression changes in real industrial operation scenarios. This provides reliable visual perception technology support for the safe collaboration, emotional interaction, and intelligent services between industrial humanoid robots and operators, promoting the engineering implementation process of bionic intelligence in industrial humanoid robots.
Under complex industrial conditions such as illumination variation, partial occlusion, and dynamic tool-workpiece contact, industrial humanoid robots require visuoperception-driven online trajectory generation for high-precision embodied manipulation. This paper proposes a real-time vision-guided online trajectory planning framework based on dynamic reverse look-ahead and piecewise NURBS reconstruction, which enables continuous trajectory updating without predefining complete reference paths. An eye-in-hand camera is adopted for real-time image acquisition, and trajectory points are efficiently extracted through ROI discretization and reverse path look-ahead. A quadratic Taylor expansion-based interpolation scheme is further developed to achieve high-precision real-time interpolation with a 2ms cycle on a real-time Linux and IgH-EtherCAT platform. Tracking experiments are conducted on three types of trajectories (line, random B-spline curve, and circular arc), with each trajectory repeated five times under different initial conditions. The proposed method achieves smooth and continuous online tracking, with an average tracking error below 0.5mm and a maximum error below 1.0mm. Moreover, ablation results show that the proposed reverse look-ahead strategy reduces the mean tracking error by approximately 30% and suppresses error oscillations compared with the baseline without look-ahead. The planning loop maintains a stable millisecond-level computational cost, validating its real-time capability for closed-loop visuoperception-driven manipulation. The experimental results validate the effectiveness of the proposed framework on a 6-axis industrial manipulator for real-time vision-guided path tracking. The current findings are most directly applicable to industrial path-following tasks such as welding, grinding, and spraying, while extension to humanoid robots should be regarded as a potential future direction that requires additional validation under whole-body balance and task coordination constraints.
With the advancement of robotics, humanoid robots have shown great application potential across various fields. This study focuses on autonomous grasping based on machine vision for humanoid robots, aiming to enhance their grasping adaptability and human-like movement capabilities in natural environments. In the machine vision module, a Realsense-D435 depth camera is adopted to collect object point cloud data, and the Iterative Closest Point (ICP) algorithm is used for point cloud registration to estimate object poses. The Denavit-Hartenberg (D-H) method models the robot's head, realizing coordinate transformation from the camera frame to the robot frame. For motion planning, referencing human arm grasping rules, the process is divided into nine basic movements, with tailored grasping poses for different objects to improve success rates. Remaining key points are autonomously calculated based on vision-derived grasping and placement points, using spatial arcs as trajectories. Matlab simulations verified the rationality of the end-effector and joint trajectories, followed by physical grasping experiments. Results show the robot can quickly and accurately recognize/locate objects in natural environments, completing grasping and carrying tasks with excellent performance: the grasping success rate reaches 89% for water bottles, 87% for small bottles, 87% for oranges, and 80% for bananas, with corresponding handling success rates of 87%, 82%, 82%, and 80%. The method ensures human-like movements while maintaining high adaptability, promoting the application of humanoid robots in daily life.
This paper proposes an industrial humanoid robot system that integrates a new type of bionic joint mechanism design and a lightweight visual perception decision-making algorithm to address the two core scientific issues of insufficient dynamic response of humanoid robot joints in sports training scenarios and the imbalance between real-time and accuracy of visual guidance interaction. This provides an optimization solution for the human-machine collaboration problem of training assistive humanoid robots. At the hardware level, this study presents an innovative design of a 7-degree-of-freedom upper limb bionic joint based on harmonic reducers and frameless torque motors. By optimizing structural parameters to solve the coupling contradiction between joint torque and response speed, the maximum output torque of the shoulder joint reached 42.3N & sdot; m, and the step response time was compressed to 62ms. At the algorithm level, a lightweight LSTM fusion pipeline was proposed to break through the accuracy-speed trade-off bottleneck of traditional visual algorithms in motion posture recognition and prediction. On the self-built SportsPace-2025 dataset, a motion recognition accuracy of 92.4% and a 3D posture estimation error of 48.7mm were achieved, with an average end-to-end delay of 168ms, meeting real-time interaction requirements. User experiments have shown that after two weeks of training with the system, the standardized score of the badminton swing motion of the humanoid robot subjects increased from 5.2 to 7.8 (p<0.01), significantly better than the control group. The research has verified the effectiveness of the proposed structural design and algorithm framework in enhancing the training assistance capability of humanoid robots and the naturalness of human-machine interaction, providing new methods for the engineering application of humanoid robots in motion scenes.
This work presents a hierarchical locomotion framework that integrates high-level, model-based step planning with reinforcement learning (RL). The Linear Inverted Pendulum (LIP) model is used to generate step timing and foot placement targets based on the robot’s current state and commanded velocity. By providing only partial guidance from the analytical model, the RL policy benefits from the predictive structure of dynamics-based planning while retaining the flexibility to overcome the limitations of simplified modeling assumptions. Compared to end-to-end learned policies, our method achieves higher sample efficiency and better generalization. The proposed approach is validated on the bipedal robot Neubot, where the learned policy enables stable walking and robust disturbance rejection. Notably, it exhibits significantly improved resilience to external perturbations compared to a fixed-step learned policy.
Taekwondo is a competitive sports event based on wisdom, which is characterized by fighting wits and courage. It is particularly important to cultivate and improve the intelligence of Taekwondo athletes in the training process. The intelligence level of athletes is the main factor determining the outcome of the competition. The intelligent learning of the athlete training course is the most important training task. Against the background of big data, Taekwondo training has made some innovations. Modern information technology has been integrated into teaching. However, in general, the application of Taekwondo training in modern information is still in the exploration stage, and training methods and strategies are still lacking, which affects Taekwondo training. This paper explores Taekwondo intelligent training and learning based on sensor technology under the background of big data and illustrates the effectiveness of this method by combining experiments. The results show that this method can achieve 93% training learning modeling accuracy and 27% training performance improvement with sensor technology.
Industrial robot incorporation into manufacturing has changed the labor market, impacting individual migration decisions (IMD). Since China is at the forefront of global industrial robot usage, the effect of automation on labor mobility has become a significant concern. This research examines industrial robot applications (IRA) and individual migration decisions (IMD) among China's floating population, seeking to understand the influence of technology developments on migration trends. A conditional logit model is employed using CMDS data and city-level robot information, modeling migration as a destination-choice process in which individuals select among competing cities based on relative utility derived from local economic conditions and automation exposure. The model examines migration choices according to individual traits, city-level controls, and IRA exposure while considering heterogeneity effects between groups and determining underlying migration mechanisms. The results show a strongly negative correlation between IRA and IMD, with a coefficient of -0.325 (p-value = 0.001), implying that greater IRA lowers migration to robot-intensive cities. The highly educated have a coefficient of -0.138 (p-value = 0.049), whereas low-educated individuals have a coefficient of -0.453 (p-value = 0.004) and are at greater risk of being displaced by automation. Migration is positively related to increased wages (coefficient = 0.578, p-value = 0.000) and negatively influenced by expensive housing (coefficient = -0.324, p-value = 0.003). The findings reflect on the intricate relationship between technology and migration patterns. Policymakers need to tackle these determinants to maximize labor market policies and offset the inequalities arising from automation.
Image captioning intends to automatically produce relevant and descriptive text for a specified image, integrating Natural Language Processing (NLP) and Computer Vision (CV) to understand visual content and express it in words. Existing image captioning methods suffer from difficulty in generating accurate and contextually rich captions, which results in captions that lack descriptive quality and alignment with visual content. The objective of this study is to develop an efficient image captioning framework capable of producing accurate and semantically rich captions from images. In this research, a hybrid Attention-reinforced transformer with contrastive learning, Serval-Frigatebird Optimization, Gaussian Error Linear Unit-Long-Short Term Memory (ArCO-SerFO-GLSTM)-based Generative Adversarial Image Captioning model is introduced for performing image captioning from a given dataset. The proposed model consists of the ArCO-SerFO generator, the Reinforcement Learning (RL Generator) and a discriminator. At first, in the ArCO-SerFO generator, the input image is passed through an image encoder to extract visual features and then fed to the caption decoder to generate a sample caption. The generated caption is compared with the ground-truth caption using contrastive loss, which improves the alignment between image features and the caption. In this case, the ArCO model is tuned exploiting Serval-Frigatebird Optimization (SerFO). The system then uses an RL generator, where an image encoder and a multi-attention mechanism guide a language decoder to generate refined captions. Then, the Reinforcement Learning (RL loss) loss updates the model based on the reward metrics. Finally, both generated captions and ground-truth captions are fed into a GELU-LSTM discriminator, which distinguishes real captions from generated captions. The GELU-LSTM is developed by incorporating a GELU into an LSTM. The developed ArCO-SerFO-GLSTM acquired Recall-Oriented Understudy for Gisting Evaluation-L (Rouge-L) of 60.19%, Mean Average Precision (mAP) of 80.13%, Bilingual Evaluation Understudy (BLEU) of 84.23%, Metric for Evaluation of Translation with Explicit Ordering (METEOR) of 31.99%, Semantic Propositional Image Caption Evaluation (SPICE) of 25.99% and Consensus-based Image Description Evaluation (CIDEr) of 123.3 with the Flickr Image dataset.
Effective communication is essential for all individuals to live a fulfilling life, but it can be particularly challenging for those who are deaf and mute. While sign language interpreters are available to help bridge this communication gap, the high cost and scarcity of qualified interpreters make it difficult for deaf and mute individuals to rely on them for everyday interactions. To address this issue, we present a real-time sign language interpretation model that converts signs into text that is easily understandable by anyone. Our approach involves analyzing the overall body movement of sign language users and using the Arabic Sign Language (ArSL) dataset requirements to train a hybrid deep learning model with a convolutional Long Short-Term Memory (LSTM) network for gesture classification. The results of our model are promising, with an average classification accuracy of 100% on the three proposed signs from the complete ArSL word dataset. Our proposed system has the potential to significantly improve the quality of life for deaf and mute individuals by removing communication barriers and increasing accessibility in everyday situations. By providing a low-cost, reliable, and efficient solution for sign language interpretation, we aim to enhance communication between deaf and mute individuals and the hearing population, fostering greater inclusivity and understanding in society.
Accurate perception and replication of dynamic human motion represent a core challenge in advancing the autonomy and naturalistic movement of humanoid robots. This research, aligned with the goals of humanoid robotics, investigates this challenge through the domain of Wushu Sanda — a complex martial art with semantically rich and varied actions. We propose a novel action recognition framework based on the semantic feature matching of salient motion images to bridge perception and imitation. Human kinematic features are first extracted via joint distance and angle calculations to form a spatiotemporal feature map. Singular Value Decomposition (SVD) is employed for data compression and redundancy reduction, facilitating the evaluation of pixel significance within salient regions for precise semantic matching. This process feeds into an optimized Deep Convolutional Neural Network (CNN) to construct a robust recognition model for Wushu Sanda movements. Experimental validation confirms the method’s efficacy, demonstrating efficient extraction of discriminative features with a maximum Gini index of 0.10, stable recognition loss converging near 0.1, and an F1-score consistently approaching 1.0 across different sample sizes. The results substantiate the proposed method’s effectiveness and underscore its potential for enhancing motion learning and imitation capabilities in humanoid robotic systems.
In the areas of robotics, sports science, virtual reality and healthcare, there has been much interest placed on human motion generation and analysis. With precise human motion modeling, HCI may enable the computer systems to comprehend, categorize, and predict these actions. Deep learning and optical motion capture technologies have opened a window to the extraction of rich spatial-temporal features from human actions for improvement in realistic motion generation and real-time analysis. A common problem faced by modern motion analysis systems is to account for the complexity in joint relations and to maintain temporal continuity, resulting in jerky and unrealistic motions during output. Additionally, some of the models are not very generalizable, particularly because they are sensitive to speed and style changes. Most conventional techniques of anomaly detection rely on manually derived thresholds, limiting their adaptability in dynamic real-world situations. This research aims to combine optical motion capture data with spatial-temporal graph convolutional networks (ST-GCNs) to overcome these limitations and establish a unified framework for motion generation and analysis. The aim is to provide reliable, scalable, realistic motion synthesis, anomaly detection, and accurate motion recognition solutions. The work is also geared towards improving model accuracy and adaptability through an efficient preprocessing and data augmentation technique. The ST-GCN spatiotemporal method for human motion modeling uses graph representations (nodes and edges) of joint movements, while the data preprocessing stages are cleaning, skeleton normalization, and data augmentation through mirroring and time scaling. The subsequent step is to build a model using a graph-based deep learning technique for motion generation, anomaly detection, and action recognition tasks. The ST-GCN model showed a high level of realism, with KL divergence values of 0.35 for “Jump” motion synthesis, while 98% action recognition accuracy was achieved for tasks like “Run” and “Punch.” Outlier anomalies detection through a threshold of 0.7 proved to be sufficient to ‘catch’ any such instances, showcasing the capabilities of the model in both analysis and generation. These results bear testimony to the effectiveness of the framework toward resolving contemporary constraining issues in motion analysis and generation. Work Related Contributions: The unified ST-GCN architecture for combined motion analysis and generation, robust preprocessing, and augmentation for improved generalization and incorporation of the graph-based anatomical constraint for preserving the kinematic constraints. It further integrates with a generative ST-GCN extension for a realistic and smooth synthesis and with adaptive anomaly detection, replacing fixed thresholds with learned spatio-temporal patterns, therefore improving novelty, usability, and scalability.
Aiming at dynamic jumping control problem of humanoid bipedal robots, parametric mechanism design, kinematic/dynamic model, numerical solution algorithm of dynamic parameters and prototype experiment of the underactuated takeoff process of humanoid bipedal robots are proposed. The underactuated takeoff process of the robot is divided into the stance phase before takeoff (dominated by the driving joint) and the underactuated takeoff phase (passively rotating around tiptoe). The kinematic model of underactuated takeoff process is established by using the D’Alembert principle, and the kinematic criterion for the robot entering underactuated phase is derived. The motion critical conditions during takeoff process are analyzed, and trajectory planning method of the total centroid based on variable quartic polynomial interpolation is proposed. The dynamic constraints of toe rotation in underactuated phase are constructed. The dynamic analytical model of underactuated takeoff is established by using the Lagrange equation, and the closed-loop solution of joint spatial trajectories is achieved through inverse kinematics. Finally, prototype experiment is carried out to verify effectiveness of the proposed mechanism design, kinematic/dynamic modeling, trajectory planning algorithm and motion control strategy.
Humanoid robots must navigate, decide, and schedule efficiently to boost automated supply chain efficiency. Traditional rule-based techniques fail in dynamic situations, especially with task dependencies. A unique Puma Optimizer-mutated Twin-Stage Adaptive Twin-Delayed Deep Deterministic Policy Gradient (PO-TSATD3) method is used in this deep reinforcement learning (DRL) system. Training datasets imitate real-world logistics situations with dynamic impediments, many robots, and varying workloads. Data preparation cleans and normalizes for quality. The Puma Optimizer optimizes convergence and operating efficiency, while the PO-TSATD3 framework improves navigation and scheduling adaptive learning. Python simulations show considerable gains in navigation accuracy, collision reduction, and schedule optimization over conventional methods. The model’s outstanding performance metrics proved its scalability and durability in complicated situations. This research validates the application of DRL, augmented by PO-TSATD3, as a powerful solution for intelligent humanoid robot operations in future supply chain systems.
With the rapid development of Internet technology and multimedia applications, and the continuous expansion of interactive scenes of humanoid robot education, the number of digital resources of ethnic music has reached an unprecedented scale and continues to grow. This study mainly focuses on the construction of a national music retrieval system that integrates multimodal attention mechanisms and its application implementation in humanoid robot education interaction. Firstly, a method for automatic annotation of ethnic music adapted to multimodal data was studied to meet the human-computer interaction needs of robot education interaction. Second, a label conditional random field music automatic annotation method is proposed, and a multi-modal attention mechanism deep neural network model for ethnic music annotation is constructed to enhance the accurate matching of music features and retrieval requirements in human-computer interaction. Finally, integrate the constructed ethnic music retrieval system into the humanoid robot education interactive platform, conduct application analysis based on actual educational interactive scenarios, verify the effectiveness of the system in the human-computer interaction process, and complete performance evaluation. The results indicate that in the Glu module, Glu blocks perform better in ethnic music annotation tasks. The annotation results of various indicators in music hierarchy sequence modeling are superior to traditional models, ensuring the accuracy of ethnic music annotation in human-computer interaction scenarios. Compared with other algorithms, the AUC label score of this system is the highest, reaching 0.913. It can more efficiently model the mapping relationship between multimodal music features input during human-computer interaction and text retrieval requirements. It shows better performance in all evaluation indicators and can effectively support ethnic music retrieval services in humanoid robot education interaction, enhancing the immersion and fun of educational interaction.
Currently, humanoid robots face high dynamics that are difficult to adapt to human movements and complex interference problems in physical education teaching scenarios. Moreover, it is difficult to accurately match the localization requirements of humanoid robots’ anthropomorphic motion trajectories, making it difficult to meet the diverse needs of physical education teaching. Therefore, this paper proposes an optimization scheme for the positioning accuracy of humanoid robots driven by visual perception collaboration, with a clear focus on biomimetic perception fusion positioning adaptation as the core solution. Construct new principles and methods for positioning optimization to address the aforementioned scientific issues. First, the system reviews the mainstream positioning algorithms and their current application status in the field of humanoid robots, with a focus on analyzing the LANDMARC algorithm and its adaptability in humanoid robot positioning. Research has found that simulating the multi-source collaborative characteristics of biological perception systems can compensate for the shortcomings of a single perception mode. The fixed weighting factor strategy of the traditional LANDMARC algorithm is unable to adapt to the dynamic positioning requirements of humanoid robot motion, which is a key bottleneck leading to positioning errors. Based on this, an improved LANDMARC algorithm is proposed that integrates biomimetic perception principles with dynamic weighted optimization. On the one hand, drawing on the multi-sensory collaborative perception mechanism of biology, a collaborative fusion model of visual perception and biomimetic tactile perception is constructed. On the other hand, a dynamic weighting factor adaptive adjustment strategy is designed, and an SVM classifier is introduced to achieve multi-round iterative optimization, forming a new method for biomimetic collaborative perception dynamic weighting iterative optimization. Simulation experiments and real physical education teaching scenarios have shown that this scheme significantly reduces the positioning error of humanoid robots in anthropomorphic motion states and complex teaching scenarios without increasing computational complexity and application costs, effectively solving the core scientific problem of positioning under dynamic motion postures. This study provides key technical support and theoretical reference for the development and application of intelligent assistive devices in physical education teaching.
Emotion detection is also a critical part of the development of human-computer interaction and mental illness treatment, where the ability to detect and respond to emotional states in real-time can have profound impacts on user experience, as well as emotional outcomes. The conventional emotion recognition systems generally cannot perform at high accuracy, particularly in real-time processing, due to challenges in extracting and fusing suitable features from facial expressions. This work introduces ResVGGNet, a novel hybrid Deep Learning (DL) model that can address these issues by integrating the advantages of VGG16 for effective feature extraction and ResNet for residual learning for DL, for better performance in real-time emotion recognition. With the accuracy of 98.60%, precision of 98.61%, and a low False Negative Rate (FNR) of 0.23%, the ResVGGNet model offers high and stable performance with various emotional states. Experimental results verify the model's ability to simplify and thus fit perfectly in emotion-aware assistive technology. This paper describes the capacity for improving users' well-being and affect regulation through real-time, dynamic feedback.
The inverse kinematics (IK) of anthropomorphic (humanoid) upper-limb manipulators presents two main challenges: the inherent non-uniqueness of solutions and inter-frame discontinuities near singularities or joint limits. Under real-time constraints, the trade-off among end-effector accuracy, near-singularity robustness, and motion smoothness becomes critical. To address these challenges, this paper proposes an on-demand global-local hybrid solver that minimizes end-effector position error, incorporates a manipulability-based robustness term, and penalizes inter-frame joint changes under a unified cost function. The solver combines damped least squares (DLS) for efficient local refinement with gated particle swarm optimization (PSO), which is activated only when increasing frame difficulty is indicated by rising error, declining manipulability, or emerging feasibility risks. At the trajectory level, smoothness and feasibility are maintained through previous-frame initialization, early stopping, and joint-limit projection. Under a unified metric suite, the proposed method is systematically compared with Always PSO-DLS, DLS-only, classical Jacobian pseudoinverse, and quadratic programming-based IK baselines. The results show that the proposed method achieves a favorable balance among tracking accuracy, continuity, robustness, and computational cost: it preserves behavior close to always PSO-DLS while substantially reducing computation, improves difficult-frame stability and reliability relative to purely local refinement, and maintains low runtime on low-complexity frames by avoiding the fixed overhead of always-on global search. Ablation and sensitivity analyses further support the effectiveness and stability of the on-demand mechanism and its key parameter settings. Overall, the proposed framework provides a promising IK core for rehabilitation robots, upper-limb exoskeletons, and anthropomorphic manipulators in continuous tracking tasks with real-time requirements.
Skeleton-based activity recognition has emerged as an essential element in intelligent surveillance systems owing to its resilience to variations in illumination, backdrop, and appearance. Graph Convolutional Networks (GCNs) have demonstrated considerable potential in modeling human motion patterns derived from skeletal data. Nevertheless, current GCN-based methodologies frequently neglect inherent topological linkages, possess restricted temporal modeling capabilities, and do not adequately represent the functional interrelation between joints and bones. We offer an innovative approach to human action recognition that integrates intrinsic bone structure with multi-scale temporal dynamics, specifically designed for real-time surveillance applications. The model incorporates an internal topological space graph convolution module that utilizes a multi-head self-attention mechanism and a common topological structure to deduce latent contextual relationships among joints. A multi-scale temporal convolution module is concurrently developed to record both fine- and coarse-grained motion patterns across different action durations. To improve feature interaction and accurately represent the structural intricacies of human movement, the model integrates a joint-bone interaction bridge, facilitating efficient fusion and transmission of skeletal data. Assessed on the NTU-RGB+D60 and 120 datasets, the suggested technique attains state-of-the-art accuracy: 91.5% (CS) and 96.9% (CV) for NTU-RGB+D60, and 89.0% (C-Sub) and 90.8% (C-Set) for NTU-RGB+D120. The results illustrate the efficacy of the suggested method in extracting comprehensive spatiotemporal features from skeletal data, providing a dependable and scalable solution for intelligent multisensor monitoring systems.
Human Action Recognition (HAR) is essential for enabling mobile robots to interact intelligently and safely in human-centric environments. This study introduces a wireless-aware multimodal fusion framework that integrates millimeter-wave (mm-Wave) radar with vision sensing for real-time HAR, ensuring robustness even under fluctuating wireless and environmental conditions. Unlike prior multimodal HAR frameworks that assume stable connectivity, our design explicitly embeds wireless channel-state feedback into the fusion process, ensuring real-time adaptability under dynamic communication conditions. The radar modality captures motion kinematics via range-Doppler and point-cloud signatures, while the camera provides spatial and appearance cues. Latency on the prototype has been kept below 50ms, and the proper integration of these heterogeneous features is additionally safeguarded through a lightweight deep-learning pipeline using adaptive fusion weighting. All the experiment simulation Outcome results showed that the proposed system obtains recognition accuracy values of up to 92% and an F1-score of 0.91, which actually outperforms both vision-only having 84% accuracy along with radar-only with 86% accuracy baselines. The results prove that wireless-vision fusion is possible for deploying real-world mobile robots in which robustness to environmental uncertainty becomes very important. The work also establishes the framework of wireless sensing in supporting context-aware human-robot interaction.
Accurate and robust Human Action Recognition (HAR) is critical for enabling intelligent systems to understand and support human activities in dynamic environments. Inclusive physical education requires flexible systems capable of fairly assessing sports actions across individuals with diverse physical abilities. Traditional assessment methods often struggle due to limited labeled data and poor adaptability to varying conditions. In this paper, we propose a Few-Shot Multimodal Sensor Fusion Framework for adaptive sports action recognition, evaluated on the opportunity and PAMAP2 datasets, to enhance generalizability and robustness. The framework leverages a transformer-based multimodal architecture, combining Vision Transformers (ViT) to extract spatial features from video data with Temporal Convolutional Networks (TCNs) to model temporal patterns from wearable sensors. To support learning from limited labeled samples, we employ few-shot learning techniques, including Contrastive Language-Image Pre-training (CLIP)-based cross-modal alignment, MetaFormer optimization, Low-Rank Adaptation (LoRA) for parameter-efficient fine-tuning, and Task-Adaptive Pretraining (TAPT) for domain generalization. Experimental results demonstrate that the framework achieves 94.3% accuracy on opportunity and 92.8% on PAMAP2, outperforming traditional baselines. High F1-scores, Area Under the Curve of the Receiver Operating Characteristic (AUC-ROC) values above 0.95, and strong few-shot generalization confirm the effectiveness of our approach. These findings highlight the potential of multimodal sensor fusion and few-shot learning for robust, scalable, and inclusive HAR, with implications for humanoid robotics, wearable systems, and intelligent physical education environments.