Long-term robotic manipulation in open environments requires unifying multimodal understanding with reliable, geometry-aware execution. Classical robotic motion planning approaches demand extensive domain modeling and hand-crafted goal specifications, while emerging LLM/VLM pipelines propose semantically plausible yet lack feasibility guarantees and executable grounding. To address the above limitations, we propose a hierarchical multimodal planning framework that combines VLM-based multimodal perception with behavior tree (BT) planning to bridge high-level semantic reasoning and low-level execution feasibility. Our framework integrates natural language instructions with open-set visual geometry to generate object-level representations and language-conditioned prototype plans. Then, a prompting-to-compilation scheme is designed to yield BT planning with explicit controller–status pairs, conditions of force, and geometric feasibility checks. The proposed framework is validated through real-world experiments across three long-term robotic manipulation tasks, showing higher task success than VLM-only and BT-only baselines and demonstrating robust, fully autonomous execution without human intervention. The paper website can be found here.
Pneumatic artificial muscle (PAM)-actuated robots always exhibit notable compliance and satisfactory human-robot interaction performance, while also suffering from complex nonlinear behaviors such as hysteresis and creep, and are highly susceptible to external disturbances. To this end, in this paper, an adaptive controller based on the fully actuated system (FAS) approach is proposed for antagonistic PAM-actuated robotic arms to achieve precise motion control without accurate system models, while maintaining strong robustness against environmental uncertainties. First, an echo state network (ESN) is elaborately integrated to estimate the systems’ unknown dynamics in real time. Then, a super-twisting extended state observer (STESO) utilizing ESN-generated information is designed to further compensate for external disturbances. Under the high-order sliding mode control framework based on FAS approaches, the plant nonlinearities are compensated by the STESO-ESN hybrid algorithm, thereby transforming the system into a closed-loop linear form. As a result, the closed-loop dynamics are flexibly shaped through pole placement, to realize high-precision trajectory tracking. Finally, a rigorous Lyapunov-based stability analysis is provided, and a series of experiments are conducted on a self-built experimental platform to validate the effectiveness of the proposed method.
As integral components in heavy-load material handling, cranes serve critical functions across numerous industrial settings including warehouse logistics, port operations, steel plants, and automated production lines. This paper focuses on overcoming the challenges associated with load rotation control in overhead crane systems by introducing a robust control strategy grounded in Super-Twisting Sliding Mode Control (ST-SMC). To tackle this challenge, a novel rotating hook mechanism is developed, and an integrated dynamic model of the overhead crane incorporating this device is established, with particular emphasis placed on characterizing the rotational dynamics. Building upon this modeling effort, a robust super-twisting sliding mode controller is synthesized to enable the rotary hook system to precisely attain desired positional targets while actively mitigating residual oscillations resulting from flexible rope suspension. Through rigorous simulation verification, the proposed control scheme is shown to deliver excellent performance and remarkable robustness across diverse operating scenarios. When evaluated against conventional Linear Quadratic Integral (LQI) control strategies, our method demonstrates measurable improvements in both load deflection reduction and oscillation suppression. These advantages prove particularly evident during operational transitions and under external disturbances, highlighting the method's potential for practical implementation in industrial crane control systems.
Bacteria are widely distributed in the biosphere, with roles ranging from beneficial microorganisms to deadly pathogens. Accurate bacterial counting is thus crucial for safeguarding human health and establishing effective medical support systems. An atomic force microscopy (AFM) is capable of generating ultra-high-resolution images, enabling researchers to quantify the number of microscopic bacteria. However, traditional manual methods are time-consuming and labor-intensive, making precise bacterial counting difficult to implement. To address the above issues, a deep learning-based approach is introduced in this paper to achieve automated bacterial counting based on AFM images. Specifically, a novel deep neural network, named as pyramid agregation MPCount, is proposed based on density map estimation to count sixteen bacterial species, including Achromobacter xylosoxidans, Burkholderia cenocepacia, Escherichia coli, etc. The proposed model achieves the minimum quantification error via training and testing on a dataset containing 588 AFM bacterial images. Experimental results demonstrate that the proposed method is efficient in identifying and counting bacteria, serving as a reference scheme for rapid and accurate microbial counting in medical systems.
Accurate tracking of flexible needles is crucial for the success of robotic-assisted puncture procedures in minimally invasive surgery. This article presents a novel sensor fusion framework that integrates permanent magnet navigation (PMN) and X-ray imaging to enable robust and precise needle tip tracking in robotic-assisted flexible needle (RFN) systems. A magnetic flexible needle is designed by embedding a permanent magnet in a nitinol spring body, allowing wireless localization while supporting the continuous rotation required by duty-cycling steering. To address challenges such as ferromagnetic interference in PMN and occlusions in X-ray, we develop a decentralized extended Kalman filter (DEKF)-based fusion algorithm that dynamically adjusts sensor weights based on cross-covariance, ensuring reliable estimation even with partial sensor failure. Experimental validation on a self-developed RFN platform demonstrates that the proposed framework outperforms single-modality systems under both magnetic interference and visual occlusion conditions, showing improved robustness and tracking continuity. This approach offers a promising solution for image-guided interventions in complex anatomical environments, with the potential to reduce the number of puncture attempts and minimize patient discomfort.
Temporal logic (TL) provides a compositional language for the formulation of long horizon robotic tasks, but existing TL-conditioned trajectory generators can sidestep perception-to-symbol binding by encoding exact object geometry in the task graph. We introduce Vision-TL-Action, which generates action trajectories from multi-view images, a coordinate-free TL syntax graph, and the robot initial state. TL-node tokens and spatial visual tokens are fused through bidirectional cross-attention, and the resulting representation conditions a flow-matching trajectory generator. Visual tokens are augmented only with normalized image-plane locations and camera-view identifiers, while a training-only predicate-to-region objective encourages grounding to referenced objects. Consistent with prior work in this domain, we evaluate the model using Success@K, the fraction of tasks for which at least one of K sampled trajectories satisfies the TL specification. On Panda task, our model achieves 67.45
Dual-boom cranes (DBCs) are widely used in industrial applications for their high load-bearing capacity, wide motion ranges, and flexible payload manipulation. However, the intricate coupling between the two cranes poses challenges for DBCs, especially in 3-D space, where unsynchronized boom motions and positioning overshoots exacerbate safety concerns. Furthermore, existing methods lack sufficient theoretical guarantees for accurate regulation of payload position and attitude in 3-D space. To address these challenges, this article proposes a novel predefined-time synchronous controller with overshoot suppression, specifically designed for 3-D DBCs. By integrating synchronization errors, overshoot suppression mechanisms, and predefined-time control, the proposed method enhances safety, coordination, and efficiency, ensuring synchronization within the same plane and coordinated motions across different planes. Rigorous static coupling analysis and closed-loop stability analysis demonstrate the method’s ability to accurately regulate payload position and attitude, while hardware experiments validate its effectiveness and robustness under diverse operation conditions. This work provides a significant contribution toward the safe and precise operation of DBCs in 3-D environments.
The unmanned aerial manipulator (UAM) system, with its ability to actively interact with the environment, has significant potential for applications in infrastructure inspection and disaster relief. However, the majority of current research focuses on interactive scenarios near the system's equilibrium point or at low speeds, limiting the utilization of the aerial vehicle's maneuverability advantages. In addition, the unique physical characteristics of hammering tasks present challenges to both the safety and effectiveness of aerial operations. To address the aforementioned problem, this article proposes an impact mitigation strategy for aerial manipulator hammering tasks, aiming to ensure the safety of the UAM system while achieving effective hammering operation. Specifically, this strategy divides the UAM hammering task into hover, dive, and hammer phases. During the dive phase, the aerial vehicle's maneuverability is exploited for accelerated descent, while the manipulator adjusts the end effector's attitude to suit the hammering scenario requirements. The planning method within the strategy considers the desired velocity and attitude of the end effector at the moment of hammering, generating smooth trajectories for the aerial vehicle and manipulator. Furthermore, a manipulator controller with active/passive work modes is designed to mitigate collision impulses without the need for any specialized devices. During the hammering phase, external and internal impulse modeling and analysis are conducted, resulting in a theoretical assurance of the maximum pre-hammering end effector velocity for system safety. Multiple sets of experimental data demonstrate that the proposed strategy performs well in hammering tasks on surfaces with inclinations and materials of different hardness, effectively achieving hammering with the UAM.
In the field of learning from demonstration (LfD), enabling robots to generalize learned manipulation skills to novel scenarios for long-horizon tasks remains challenging. Specifically, it is still difficult for robots to adapt the learned skills to new environments with different task and motion requirements, especially in long-horizon, multistage scenarios with intricate constraints. This article proposes a novel hierarchical framework, called BT-TL-DMPs, that integrates behavior tree (BT), temporal logic (TL), and dynamical movement primitives (DMPs) to address this problem. Within this framework, signal temporal logic (STL) is employed to formally specify complex, long-horizon task requirements and constraints. These STL specifications are systematically transformed to generate reactive and modular BTs for high-level decision-making task structure. An STL-constrained DMP optimization method is proposed to optimize the DMP forcing term, allowing the learned motion primitives to adapt flexibly while satisfying intricate spatiotemporal requirements and, crucially, preserving the essential dynamics learned from demonstrations. The framework is validated through simulations demonstrating generalization capabilities under various STL constraints and real-world experiments on several long-horizon robotic manipulation tasks, where it significantly improves the success rate from 0.25 to 0.85, compared with the baseline. The results demonstrate that the proposed framework effectively bridges the symbolic-motion gap, enabling more reliable and generalizable autonomous manipulation for complex robotic tasks.
The complex interactions between flexible needles and tissues present significant challenges in predicting the needle shape during the puncture procedure. In particular, the accurate prediction of flexible needle shape during insertion into complex multilayer tissues, especially when measurement feedback involves non-Gaussian noise, remains an open problem. In this article, we develop a novel reinforcement learning-based active modeling scheme to predict the deflection of the robotic flexible needle. First, the active modeling scheme is constructed by deriving an extended Kalman filter under the maximum correntropy criterion to enhance insensitivity to non-Gaussian noise. Subsequently, based on this scheme, the reinforcement active modeling (RAM) framework is built by incorporating reinforcement learning to compensate for the modeling residuals. Specifically, the theoretical convergence of the proposed scheme is proved by using the Banach fixed-point theorem, thereby ensuring the reliability of needle shape prediction. Finally, a series of comparative experiments is carried out on a self-built robotic flexible needle. The experimental results demonstrate the superior performance of the proposed deflection predictor. Under non-Gaussian noise conditions, the proposed RAM scheme achieves a generalization prediction error reduction of 46.4% in RMSE and over 76.1% in Var during insertion into unknown multilayer tissue.
This paper addresses the challenging problem of dynamic ball-hitting by the end-effector of a 2-degree-of- freedom manipulator equipped on the multirotor. Traditional approaches that rely on fixed hitting heights often lack adaptability to varying ball trajectories, limiting performance in dynamic scenarios. To address this, the hitting task is formulated as a Markov Decision Process (MDP) and solved using a reinforcement learning framework. A hierarchical action space combining big/small step sizes with hysteresis logic is designed to balance trajectory smoothness and learning efficiency. Furthermore, a dynamic hitting point selection strategy enables the multirotor to adaptively determine the optimal timing and position for interaction. Leveraging experience replay and target network updates, the Deep QNetwork (DQN) framework ensures stable and robust training. Simulation results show that the proposed method achieves higher success rates, smoother trajectories, and stronger stability than baseline strategies, demonstrating its potential for dynamic aerial manipulation tasks.
This article presents a novel learning-based approach for online state estimation in flapping-wing aerial vehicles (FWAVs). Leveraging low-cost magnetic, angular rate, and gravity (MARG) sensors, the proposed method effectively mitigates the adverse effects of flapping-induced oscillations that challenge conventional estimation techniques. By employing a divide-and-conquer strategy grounded in cycle-averaged aerodynamics, the framework decouples the slow-varying components from the high-frequency oscillatory components, thereby preserving critical transient behaviors while delivering a smooth internal state representation. The complete oscillatory state of FWAV can be reconstructed based on above two components, leading to substantial improvements in state estimation accuracy. Experimental validations on an avian-inspired FWAV demonstrate that the estimator enhances accuracy and smoothness, even under complex aerodynamic disturbances. These encouraging results highlight the potential of learning algorithms to overcome issues of flapping-wing induced oscillation dynamics.
Salamander-like robots, designed to mimic the morphology and locomotion of real salamanders, exhibit impressive agility and adaptability, making them suitable for autonomous tasks such as environmental monitoring and disaster relief. To perform these tasks effectively in complex environments, robust visual navigation is essential. However, existing fuzzy-logic methods scale poorly and require extensive manual tuning, while approaches developed for other bio-inspired robots often rely on rigid motion models or large labeled datasets, limiting their applicability to salamander-like platforms. To address these challenges, this paper presents a hierarchical navigation framework that integrates deep reinforcement learning (DRL) with biologically inspired control. A DRL-based policy learns to map stacked depth observations and goal information to a compact control signal, which modulates a low-level motion controller to generate coordinated gaits. The proposed method enables adaptive, end-toend navigation without relying on explicit models or handcrafted controllers. Simulation results in cluttered environments show that the proposed method consistently outperforms the baseline method, demonstrating more stable performance and stronger generalization, thereby underscoring its potential to advance vision-based autonomy in salamander-like robots.
In this paper, a new environment-adaptive navigation strategy is proposed for multirotor swarm, which guarantees flexible, efficient, and safe navigation through deep learning. The proposed navigation strategy consists of two parts. In the path-finding stage, the maximum gap path planning algorithm (MGPP) is designed to search for a feasible path for the geometric center of the swarm. Building upon this, a distribution planning method is proposed for the multirotor swarm, enabling environment-adaptive flight even in complex, obstacle-dense environments. Specifically, a new swarm distribution planning framework is proposed by designing two deep neural networks (DNNs), where the shape and orientation angle of the swarm are online generated according to real-time local scene images. In this way, scene images are continuously fed into onboard networks and online planning swarm distribution. Finally, individual priorities are assigned to each multirotor to prevent collisions among the paths generated from the swarm distribution planning results. Flexible, environment-adaptive navigation is ensured by integrating front-end path finding with back-end distribution planning, resulting in both efficient and safe flight operations. The simulation experiment results are included to show the effectiveness of the proposed strategy.
Head stabilization is crucial for 3-D snake robots to perceive their environment effectively using onboard vision sensors. This article proposes a head-orientation-stabilized gait that enables the snake robot to capture wide-field video with improved stability. Given the challenge of compensating for coupled swings in both directions, the snake robot is divided into a motion part and a head-raising part, as guided by the proposed design rules. The motion section suppresses body-induced disturbances while preserving off-road locomotion capability and the head-raising section uses a hook shape to broaden the camera view and reduce residual disturbances. To generate the motion part, the S-pedal gait, suited for uneven terrain, is reformulated into Frenet-Serret-based angle functions for localized body tuning without full-shape transmission. Furthermore, setting specific rolling angles leverages segment twist angles to decouple 3-D shape formation, where pitch and yaw joints generate floating and grounding arcs, respectively, enabling independent joint modulation and enhanced swing suppression. A series of experiments validates the proposed gait, with results showing satisfactory head stabilization, wider camera coverage, and improved adaptability to challenging terrain compared to existing methods.
Autonomous ultrasound scanning requires a robotic manipulator to maintain stable force contact while moving on deformable surfaces. Traditional reinforcement learning (RL) often results in jerky, unnatural motions, while standard imitation learning requires extensive expert datasets. This paper presents a novel framework for autonomous ultrasound scanning that integrates adversarial motion priors (AMP) with a multimodal state representation. By fusing U-Net extracted visual features with proprioceptive and force feedback, our proposed hybrid architecture enables the learning of scanning policies that are both efficient and human-like. The system utilizes a discriminator trained on expert demonstrations to reward motion style, while simultaneously optimizing for task-specific goals such as stable force contact maintenance. Simulation results in a simulated robotic ultrasound environment demonstrate that our multimodal adversarial reinforcement learning approach significantly improves sample efficiency, motion smoothness, and contact stability compared to baseline PPO methods.
Active perception in vision-based robotic manipulation aims to move the camera toward more informative observation viewpoints, thereby providing high-quality perceptual inputs for downstream tasks. Most existing active perception methods rely on iterative optimization, leading to high time and motion costs, and are tightly coupled with task-specific objectives, which limits their transferability. In this paper, we propose a general one-shot multimodal active perception framework for robotic manipulation. The framework enables direct inference of optimal viewpoints and comprises a data collection pipeline and an optimal viewpoint prediction network. Specifically, the framework decouples viewpoint quality evaluation from the overall architecture, supporting heterogeneous task requirements. Optimal viewpoints are defined through systematic sampling and evaluation of candidate viewpoints, after which large-scale training datasets are constructed via domain randomization. Moreover, a multimodal optimal viewpoint prediction network is developed, leveraging cross-attention to align and fuse multimodal features and directly predict camera pose adjustments. The proposed framework is instantiated in robotic grasping under viewpoint-constrained environments. Experimental results demonstrate that active perception guided by the framework significantly improves grasp success rates. Notably, real-world evaluations achieve nearly double the grasp success rate and enable seamless sim-to-real transfer without additional fine-tuning, demonstrating the effectiveness of the proposed framework.
Risk-aware navigation in unknown environments is a fundamental challenge for autonomous vehicles operating in complex urban systems. To address this issue, this paper presents a differentiable optimization layered safety-critical control method based on conformal prediction. First, to handle uncertainties arising from sensor noise, the conformal prediction method is employed to generate risk-aware obstacle ellipsoids around an elliptical-shaped robot. Second, two nested differentiable optimization layers are introduced to build the control barrier functions for obstacle avoidance and feasibility guarantee, respectively. Then, a quadratic program based safety-critical control law is proposed to integrate the above control barrier function constraints as well as input constraints. In the end, the effectiveness of the proposed framework is demonstrated through numerical simulations.
Gaussian process regression (GPR) models are becoming increasingly tightly integrated into robotic systems, particularly in the context of robot model predictive control (MPC) operating in complex environments. Because data generated by robots are typically collected online and exhibits heteroscedasticity (i.e., the noise variance depends on the input), traditional GPR may not be suitable. Thus, an incremental heteroscedastic GPR (IHGPR) method is proposed, which takes advantage of incremental sparse spectrum GPR (I-SSGPR) and the framework of improved most likely heteroscedastic GPR (improved MLHGPR). The predictive distribution is not only in an explicit form but also differentiable, rendering a plug-and-play solution for optimization-based control. The efficacy of the proposed approach is demonstrated through a series of empirical evaluation experiments, highlighting its time efficiency and capacity to accurately capture heteroscedastic signals. To illustrate its versatile applicability in robotic system control, we integrate IHGPR into a robust MPC (RMPC) method to online fit the state-and input-dependent heteroscedastic stochastic disturbances and present the applicability and efficiency through simulation results.
Cable-suspended aerial transportation systems are employed extensively across various industries. The capability to flexibly adjust the relative position between the multirotor and the payload has spurred growing interest in the system equipped with variable-length cable, promising broader application potential. Compared to systems with fixed-length cables, introducing the variable-length cable adds a new degree of freedom. However, it also results in increased nonlinearity and more complex dynamic coupling among the multirotor, the cable and the payload, posing significant challenges in control design. This paper introduces a backstepping control strategy tailored for aerial transportation systems with variable-length cable, designed to precisely track the payload trajectory while dynamically adjusting cable length. Then, a cable length generator has been developed that achieves online optimization of the cable length while satisfying state constraints, thus balancing the multirotor's motion and cable length changes without the need for manual trajectory planning. The asymptotic stability of the closed-loop system is guaranteed through Lyapunov techniques and the growth restriction condition. Finally, simulation results confirm the efficacy of the proposed method in managing trajectory tracking and cable length adjustments effectively.