
This study proposes a multimodal robotic task planning framework based on large language models (LLMs) that utilizes both visual and force-derived physical feedback. The system integrates multimodal inputs, including camera images and force/torque (F/T) sensor data, to interpret the visual and physical properties of objects and generate a task sequence to execute given commands. The integration of vision and force data allows the system to compensate for the limitations of each modality, leading to more reliable decision-making. The framework consists of three components: Extractor, Planner, and Sub-planner. According to their assigned roles, each agent automatically translates high-level commands given in natural language into executable robot motion plans, enabling the robot to perform the required sequence of actions. The system combines visual and force data to handle tasks that are difficult or infeasible with a single modality. Furthermore, it supports sequential and condition-based task planning through the collaboration of multiple LLM agents. Experimental results in various task scenarios show that the proposed framework consistently improves overall task success rates compared with unimodal settings with different LLMs and achieves a higher success rate compared to using only visual or force data.
Industry 5.0 promotes collaborative manufacturing environments where humans and robots work together, leveraging their complementary strengths to improve productivity and safety. In such settings, trust is fundamental to ensuring seamless and effective interaction. As intelligent robots take on more complex roles, systems must be capable of adapting to changes in human trust to maintain safe and efficient operations. However, most existing trust assessment methods rely on post hoc questionnaires, which do not capture observable behavioral cues during interaction that influence human decision-making, workload, and task performance. This study addresses this limitation by introducing a multi-modal, data-driven machine learning model for classifying behavioral patterns associated with distinct trust states under controlled trust-inducing conditions. The framework integrates facial expression features extracted using a CNN model with body motion indicators processed through kinematic models. The model was evaluated in a chemical industry scenario, where a robotic manipulator supports operators in handing and mixing hazardous materials. It achieved 88.40
This study examines the impact of the social robot LIAN on the development of socio-emotional, attentional, and attitudinal skills in neurodivergent and neurotypical students, with the goal of evaluating its potential as a teaching support tool in inclusive educational environments. LIAN was designed for pedagogical use and integrates verbal and nonverbal communication that enables structured and engaging interactions with children. In this study, the robot reinforced teacher instructions and supported student participation during classroom activities. Although LIAN includes an artificial intelligence module capable of generating dynamic responses, this study used a controlled, pre-programmed configuration instead to ensure consistency and reduce variability between sessions. The case study involved students with a confirmed diagnosis of Autism Spectrum Disorder (ASD) and a neurotypical comparison group. Six variables were assessed, and teachers were interviewed using models of technological acceptance and pedagogical knowledge. Overall, the results indicate that LIAN provided specific benefits for students with ASD in several key areas.
This study presents a cost-effective, modular, and contamination-tolerant alternative to conventional servo valves for proportional flow control in hydraulic service robots. Although servo valves provide high precision and bandwidth, their mechanical complexity, cost, and sensitivity to contamination limit their suitability for miniaturized and budget-constrained robotic systems. A hybrid flow control architecture is proposed that combines multiple ON/OFF solenoid valves with binary-weighted orifice configurations and pulse width modulation (PWM) for fine resolution. In contrast to conventional digital hydraulic schemes that apply high-frequency switching signals to every valve in the array, the proposed selective-PWM strategy restricts duty-cycle modulation exclusively to the smallest-flow valve, while the larger valves are operated only as discrete static stages. This selective modulation reduces switching-cycle accumulation and local switching transients in the larger valves while preserving the effective flow resolution of a fully switched array, providing a cost-effective, modular, and contamination-tolerant alternative for moderate-bandwidth hydraulic flow control. Flow-level experimental characterization was conducted under both low-pressure (water-based) and high-pressure (100 bar hydraulic oil) conditions. The results demonstrate a linear relationship between flow rate, valve combinations, and PWM duty cycle within the 10–80
This study explores and experiments with the development of a social robot as a support tool for the teacher with respect to the learning of communication and socioemotional skills of children with Special Educational Needs (SEN). The objectives include its evaluation in a school context, being a support tool in an activity dictated by the teacher, to enhance its design through the analysis of this first human-robot interaction. Considering factors such as emotional expression, receptivity and expectations of children. This approach will allow identifying areas of improvement in its functional design, ensuring that the robot adapts to specific needs and facilitates a more effective and enriching interaction in educational environments.
The difficulty of robotic manipulation often depends strongly on object pose, making the reorientation of horizontally placed objects into vertical poses an important preparatory step for subsequent tasks. Conventional pivoting methods often rely on firm grasps, carefully designed trajectories, or forceful interactions with external surfaces, which can reduce robustness and increase task complexity. This paper presents a reinforcement learning–based Half-grasping strategy that achieves controlled rotational slip for efficient object pivoting by adaptively modulating the grasping force online, primarily leveraging gravity while strategically utilizing essential, minimal ground contact. The policy was trained entirely in simulation with domain randomization over object mass and CoM variations and was successfully transferred to a real robot without task-specific fine-tuning. Quantitative experiments under varying mass, CoM, friction, and object-geometry conditions showed that the proposed method generally outperformed the baselines in mass and CoM variations, particularly in terms of success rate, while maintaining robust performance across the evaluated conditions. Across the 22 non-overlapping test conditions summarized in the quantitative experiments, the proposed method achieved an overall success rate of 86.7 % . These results support the effectiveness of grasp-force modulation based on Half-grasping for robust pivoting-based reorientation in real-world settings.
Intersection over union (IoU) is the most popular metric for evaluating accuracy in robot manipulation. However, classic IoU loss is insufficient to maximize the value of metric in robot manipulation. There are many IoU-based loss which are designed for deep learning, incorporating penalty terms to improve upon the classic IoU loss. Moreover, these methods still encounter suboptimal performance in neuro-symbolic robot manipulation (NSMR) which incorporate both symbolic reasoning and deep learning. Therefore, to accelerate convergence speed and improve the manipulation accuracy in NSMR, this paper presented distance–area–IoU fusion (DAIoU_F) loss which combines information from both line and area dimensions of the bounding box corners. Three diverse language-guided robot manipulation datasets each containing different object and object rational language instruction have been built in this paper. DAIoU has been integrated into the latest language-guided robot manipulation neuro-symbolic operation architecture and verified on three diverse datasets. This paper achieves improvements in both the IoU metric and convergence speed within the same training period. The dataset, checkpoint and code are available in https://github.com/cher0000/language-guided-robot- .
The rapid rise of industrial automation has accelerated the deployment of autonomous mobile robots for material handling, inspection, and collaborative operations. Effective performance in dynamic environments demands path-planning algorithms that ensure safety, energy efficiency, and adaptability—balancing global optimality with adaptive reactivity. No single algorithmic solution fully satisfies these requirements, necessitating hybrid frameworks that integrate the global optimization capability of metaheuristics with the adaptive learning of reinforcement learning. To address these challenges, we introduce RL-PFWOA, a novel hybrid hierarchical framework featuring bidirectional feedback between a metaheuristic core and a reinforcement learning agent. Global exploration is achieved through a hybrid Pufferfish optimization (PFO) and whale optimization algorithm (WOA) strategy, where the WOA coefficient A dynamically alternates between encircling and spiral exploitation phases. A DDPG agent performs path refinement and adaptively modulates exploration via hybrid weights w_PFO and w_WOA . A Bayesian weighting mechanism further balances path length, energy consumption, and traversal time, ensuring adaptive multi-objective optimization and preventing premature convergence. Benchmarking on ten standard functions (Sphere, Rastrigin, Schwefel, etc.) demonstrates superior convergence and minimal cost values compared to metaheuristic (PSO, GA, ABC, WOA) and reinforcement learning (PPO, SAC) baselines. MATLAB/Simulink 2023a simulations validate practical performance: RL-PFWOA yields the shortest paths (55.0 m static, 58.1 m dynamic), lowest energy use (9.21 J static, 9.87 J dynamic), and minimal obstacle collisions (1.5–2.1 SD = ± 0.35 ), a 56
In recent years, cooperative robot path planning has gained significant attention because of its wide range of applications, including load transportation and firefighting operations. However, existing methods based on classical control or hybrid strategies that combine control with deep reinforcement learning (DRL) often face challenges in generalizing and adapting to dynamic environments or unexpected situations. To address these issues, this work presents a fully DRL-based safe path planning approach for leader–follower robotic systems. In particular, we propose a modified Multi-Agent Twin-Delayed Deep Deterministic Policy Gradient (M-MATD3) algorithm, specifically designed to mitigate common issues such as overestimation bias and high variance observed in standard MATD3. Extensive simulations under a centralized training and decentralized execution (CTDE) framework show that the proposed M-MATD3 algorithm demonstrates strong performance, lowering collision rates to 8.15
Prosthetic knee joints play a critical role in restoring gait function in individuals with lower-limb amputation; however, conventional evaluation methods involving human participants are constrained by limitations related to safety, ethics, cost, and experimental repeatability. To address these challenges, this study proposes a robotic gait simulation platform for systematic evaluation of prosthetic knee kinematics and control performance. The proposed 3-DOF gait simulator generates hip trajectory-driven motion through synchronized horizontal translation, vertical displacement, and hip flexion–extension rotation. Validation experiments demonstrated that the simulator reproduced reference hip joint trajectories with high accuracy. Using the proposed platform, both a passive prosthetic knee (Ottobock 3R60) and a microprocessor-controlled prosthetic knee (MPK) were experimentally evaluated. The passive prosthetic knee exhibited simulated knee joint trajectories showing high agreement with experimentally measured transfemoral amputee knee kinematics, and the MPK demonstrated highly repeatable and stable performance under repeated operating conditions. Threshold-based control verification confirmed reliable operation of the control logic under various input conditions, and cyclic loading tests verified structural reliability under repetitive operating conditions. Furthermore, knee joint torque analysis revealed consistent cyclic torque generation patterns during simulated gait. These results demonstrate that the proposed simulator can reliably generate hip trajectory-driven motion based on predefined reference trajectories and serve as a repeatable platform for quantitative preclinical assessment of prosthetic knee kinematics, control performance, and mechanical reliability without requiring human participants.
Unmanned aerial vehicles (UAVs) are increasingly employed in applications such as surveillance, disaster management, traffic monitoring, and precision agriculture. However, real-time object detection platforms remain a challenging task due to limited onboard computational resources, stringent latency requirements, dynamic network conditions, and complex environmental variations including illumination changes, occlusion, fast motion, and low-resolution targets. Existing UAV-based detection systems often rely on static edge or cloud processing strategies, which limits their adaptability and degrades performance in dynamically changing operational scenarios. To address these challenges, this paper proposes an Adaptive Reinforced Detection and Tracking Network (ARDT-Net), which integrates a lightweight YOLOv9-Lite detector, Soft Actor-Critic (SAC)-based adaptive offloading, and reinforcement learning object tracking (RLOT) for robust edge–cloud UAV in complex environments. The framework employs YOLOv9-Lite as a lightweight edge-side detector to enable low-latency inference on resource-constrained UAV platforms. To efficiently balance computational load and detection accuracy, a SAC-based reinforcement learning mechanism is introduced to dynamically control task offloading decisions between the UAV edge and the cloud, considering workload complexity, network conditions, and available system resources. Furthermore, a RLOT module is integrated at the cloud level to refine detection results and maintain temporal consistency under challenging conditions. The proposed framework also incorporates efficient feature extraction and multi-scale feature aggregation to enhance robustness while preserving computational efficiency. Extensive experiments conducted on the UAVDT, VISDRONE, AU-AIR, and DRONEVEHICLE datasets, which includes diverse lighting conditions, viewpoints, and low-resolution targets, demonstrate the effectiveness of the proposed approach. The ARDT-Net framework consistently achieves superior precision and success rates compared to State-Of-The-Art (SOTA) UAV object detection methods. Ablation studies further confirm that the SAC-driven adaptive offloading strategy and the RLOT module achieve robust, low-latency, and accurate UAV object detection in real-world scenarios.
This study investigates and enhances the accuracy of real object detection in service robots, particularly in environments where printed images or advertisements may interfere with the detection of real-world objects. A hybrid approach is proposed, combining YOLO models for object detection, depth estimation techniques for spatial understanding, and edge-texture analysis for distinguishing real objects from images. The COCO dataset, containing 2D images, and ShapeNet, providing 3D models, were used to construct a custom evaluation dataset for benchmarking the proposed hybrid approach using pretrained YOLOv8 and MiDaS models. Experimental results show that the integration of depth estimation and edge‑texture analysis achieved an accuracy of 86
In search and rescue operations, there is a period known as the “golden time” during which the probability of finding the target alive is highest. The objective of this work is to propose a new search algorithm for unmanned aerial vehicles (UAVs) with a focus on improving the detection probability and execution time. We approach this problem by first modeling target dynamics as a Markov process and the detection likelihood as a function of image quality and the observer’s vision. We then employ Bayesian theory to derive a fitness function representing the probability distribution of the target’s location over the search area. Finally, we introduce a new algorithm named polar coordinate-based differential evolution (PDE) to generate a UAV search path that maximizes this fitness function. The PDE algorithm utilizes polar coordinates to incorporate kinematic constraints and maneuver properties of the UAV, allowing for better exploration of the solution space. A series of simulations and comparative analyses have been conducted to evaluate the performance of the proposed algorithm. Experiments involving a real UAV have also been conducted. Results demonstrate that the PDE algorithm outperforms state-of-the-art algorithms in terms of detection probability and execution time across diverse search scenarios while remaining practical for real-world applications. The source code of the algorithm is available at https://github.com/thuhangkhuat/PDE_target_search .
Legged robots often face a trade-off between stability and energy efficiency, as postures that improve stability typically increase energy consumption, particularly in unstructured environments such as disaster response or exploration sites. While additional appendages can introduce a synergy between stability and efficiency, they also increase hardware and control complexity. In this study, we show that the strategic use of intrinsic kinematic redundancy through posture modulation can provide a stability–efficiency synergy. We design a stability-enhanced gait for a redundant quadruped robot that independently modulates body height z_s and foot orientation n_y under fixed foothold conditions, and experimentally evaluate its energetic performance. As evaluation metrics, we use the minimum normalized energy stability margin (NESM) over a gait cycle to quantify static stability and the cost of transport (CoT) to assess energy efficiency. Under quasi-static, no-slip, slope-walking conditions, experimental results reveal a strong negative correlation between the minimum NESM and CoT ( ρ =-0.86, p<0.001 ). In particular, configurations with a lower body height and a foot orientation slightly exceeding the slope angle increased the maximum slope from 10^∘ to 20^∘ (100
This paper presents a closed-loop, proprioception-driven adaptive locomotion strategy for a quadruped robot with an actuated spinal joint, aiming to enhance locomotion stability, terrain adaptability, and energy efficiency on complex rigid terrains. Unlike conventional rigid-body quadruped platforms, the proposed system incorporates an active spine-limb coupling mechanism. We establish a full-body kinematic model incorporating the spinal degree of freedom and develop a hierarchical central pattern generator (CPG) network based on Hopf oscillators to coordinate rhythmic spinal and limb motions. To address perception limitations in complex environments, we design a heuristic, threshold-based proprioceptive terrain classification framework that fuses kinematic data with contact states to classify terrain features exclusively for rigid terrain scenarios. A bio-inspired reflex mechanism is synthesized with the CPG to dynamically regulate joint equilibrium positions, body posture, and foot trajectories, ensuring adaptive stability on slopes and rugged terrains. Both simulations and prototype experiments validate the effectiveness of the proposed strategy, demonstrating significant improvements in stability, adaptability, and energy efficiency.
This paper presents an on-site kinematic calibration framework for robotic manipulators using a single-laser distance sensor. The proposed method models the discrepancy between nominal and measured laser distances as a function of the modified Denavit–Hartenberg (MDH) parameters, which are identified through iterative least-squares estimation using an analytically derived identification Jacobian matrix. To ensure efficient and high-quality data collection, a nonlinear programming (NLP)-based configuration sampling strategy is developed that autonomously generates informative measurement poses while strictly satisfying operational constraints, including sensor range, incidence angle, and collision avoidance. The proposed sampling procedure generates 480 configurations across two orthogonal measurement planes in approximately 17 s , representing a substantial reduction in effort compared to manual pose teaching. Experimental results demonstrate a 76.87% reduction in reprojection root mean square error (RMSE), decreasing from 1.590 mm to 0.368 mm , with consistent performance across training and validation sets. Independent validation using a laser tracker confirmed sub-millimeter end-effector position accuracy, with a combined RMSE of 0.679 mm over 10 configurations. Additional experiments on four target surface materials revealed that surface flatness quality, rather than optical finish, is the dominant factor governing measurement accuracy during robot motion. These results demonstrate the practical effectiveness of the proposed framework for on-site robot accuracy enhancement without reliance on high-precision metrology equipment. The proposed framework achieves sub-millimeter end-effector accuracy using a single-laser distance sensor, offering a substantially lower-cost alternative to manufacturer recalibration services that typically involve shipping costs, service fees, and multiple days of robot downtime.
In simultaneous localization and mapping (SLAM), recognizing previously revisited locations, a task known as loop closure detection, is crucial for correcting accumulated drift and ensuring reliable navigation, especially in large-scale or multi-robot mapping scenarios with numerous keyframes. In LiDAR-based loop closure detection, a bird’s-eye-view (BEV) projection combined with a k-d tree is a widely used approach due to its efficiency. However, this combination can lead to significant information loss when compressing a 3D point cloud into a compact fixed-length global descriptor, which may exclude true loop candidates and result in missed loop closures. To address this limitation, we propose DZLoop, a novel LiDAR-based loop closure detection method that leverages multiple structural density images and Zernike moments. By combining these components, DZLoop generates a robust global descriptor that effectively captures diverse structural characteristics of the environment. Unlike conventional methods that obtain rotation invariance by compressing rows or columns of the BEV image, Zernike moments inherently provide image-level rotation invariance. Experimental evaluations on public datasets, including KITTI and HeLiPR, demonstrate that DZLoop outperforms existing methods such as M2DP, PALM, ScanContext, and NDD, achieving more reliable loop detection performance.
When applied to mobile robot path planning, the traditional Rapidly-exploring Random Tree (RRT) algorithm frequently suffers from inadequate sampling efficiency, sluggish planning speed, and suboptimal path quality. To address these inherent challenges, an optimized sampling-based RRT algorithm (OS-RRT) is proposed in this paper. Specifically, the OS-RRT algorithm introduces four key enhancements: (1) increasing the target-biased sampling probability to accelerate tree expansion and boost search efficiency; (2) deploying temporary candidate nodes near obstacles to minimize the waste of sampling points and ensure effective tree growth; (3) dynamically constraining the sampling space to significantly curtail blind exploration in invalid regions; and (4) implementing a dual-segment path optimization strategy to further elevate both planning speed and final path smoothness. The effectiveness of the above four methods is verified by experimental comparison. The superiority of the performance of the OS-RRT algorithm is verified by comparing it with the existing mature path planning algorithms (RRT algorithm, ACO algorithm and Bias-RRT algorithm) in maps of different complexity. The experimental results show that in three different environments, compared with RRT algorithm, the average path length of the OS-RRT algorithm is shortened by 20.00
Enabling intuitive, low-level robot control through natural spoken language is essential for deploying service robots in everyday environments. However, verbal commands are inherently variable and often ambiguous, making it difficult for robots to consistently interpret and execute them without extensive handcrafted rules or predefined templates. In this paper, we present the directive language model (DLM), a novel speech-to-trajectory framework that directly maps unconstrained verbal instructions to low-level robot motion trajectories. DLM is trained using behavior cloning (BC) on demonstrations in simulation, where human participants issue spoken guidance and control the robot’s motion accordingly. To improve generalization across phrasings, we apply semantic augmentation via GPT-4, generating diverse paraphrases that share the same trajectory labels. A text-conditioned diffusion policy is then employed to generate smooth and flexible trajectories that align with user intent. Unlike large language model (LLM)-based approaches that require prompt engineering and often yield unpredictable responses, DLM ensures consistent motion generation with efficient inference suitable for onboard deployment. Our results, validated in both simulation and on a quadruped robot, show that DLM robustly interprets a wide range of user commands, including paraphrased, truncated, and implicit instructions, without relying on perception or symbolic planning. These capabilities make DLM well-suited for natural, speech-based control of service robots interacting with untrained users.
Scene rearrangement is a crucial capability for household robotic assistants, requiring an embodied agent to restore objects to a previously recorded configuration after external modifications. This work addresses the core challenges of scene understanding and high-level task strategy , which depend on an agent’s ability to accurately identify objects, ascertain their states, and detect discrepancies from a target layout. Despite recent progress, a gap remains for methods that can operate without privileged information, such as complete scene geometry or ground truth object states. This paper presents a complete, end-to-end framework designed for such realistic constraints in complex kitchen environments. The proposed methodology integrates deep learning-based perception with spatio-temporal analysis of egocentric RGB-D data. This enables the system to detect object relocations, identify intrinsic state changes, infer necessary tool-based actions, and plan an optimal execution sequence for the entire task. Evaluated in the AI2-THOR simulation environment, the system achieves a 93.6