
In this article, we propose ContactFlowFilter, a novel approach for contact estimation that determines the number of active contacts, their locations, and the associated force vectors. We consider a robot manipulator equipped with joint torque sensors and a base force/torque sensor; such a proprioceptive-only setup presents a challenging inverse problem characterized by (i) multi-modality from non-unique solutions, (ii) temporal dependencies in state evolution, and (iii) discrete contact events involving contact addition/removal. Existing approaches, such as particle filters (PFs) and learning-based methods, often struggle to address these challenges simultaneously. To overcome this, ContactFlowFilter integrates a generative model into the PF framework. The recursive structure of the PF captures temporal dependencies, while the flow matching model serves as an auxiliary proposal distribution to track multi-modal beliefs and detect contact events. Extensive validation in both simulation and real-world experiments confirms the effectiveness of the proposed method. In real-world settings, ContactFlowFilter achieves the lowest dissimilarity from the target distribution, with errors of only 0.36 cm and 0.88 cm for single- and dual-contact scenarios, respectively, while operating at 0.69 ms and 2.69 ms.
This paper presents an innovative guidance and control method for a space manipulator that actively maneuvers its base to hold a tumbling object stationary relative to the base during proximity operations. This simplifies motion planning, reduces collision risk, and enables impact-free capture using single- or multi-arm configurations, with or without predefined grasping interfaces. The proposed approach involves two phases. In the pre-grasping phase, a synchronization control aligns the target's center of mass (CoM) with the system's instantaneous center of rotation (ICR), virtually rigidizing the target–servicer system to keep the object stationary for reliable grasping. In the post grasping phase, a control strategy based on Hamiltonian dynamics and optimal control stabilizes the combined system by minimizing time, fuel, or energy, while respecting base torque limits and multiple grasping constraints. For structured targets, applied forces and torques stay within their bounds; for unstructured ones, friction constraints prevent slipping. A case study demonstrates the method's effectiveness in enabling both the subsequent grasping and de-tumbling phases.
Robot calibration is a fundamental technique to ensure precision and performance across diverse robotic systems and applications. This study proposes a unified calibration framework based on optimization over Riemannian product manifolds. The framework is aimed at overcoming the limitations of existing methods that require problem-specific formulations. The proposed approach preserves intrinsic geometric structures by performing optimization directly on manifolds without applying scalar approximations. This enhances both theoretical rigor and computational efficiency. The key contributions include the following: (1) a general formulation of calibration problems as manifold mappings with adapted Jacobian constructions, (2) a robust iterative optimization method with null space elimination to address ill-posed problems, and (3) an extension of Bayesian optimization to Riemannian manifolds using positive-definite kernels and Levi–Civita parallel transport for global initialization. Simulation and real-world experiments demonstrate the framework's superior accuracy, generalizability, and computational efficiency across a wide range of calibration scenarios.
Guiding vector field (GVF) methods provide effective solutions for manifold-following problems by generating smooth guidance signals that steer robots toward and navigate desired geometric manifolds. Recent advances in high-dimensional GVF design have successfully eliminated singularities that arise in classical formulations. However, robots must additionally achieve obstacle avoidance, inter-robot collision prevention, and cooperative coordination in practical applications. When vector-field composition is introduced, existing GVF-based approaches often fail to preserve the singularity-free property. To address this challenge, a truncated guiding vector field and an obstacle-avoidance vector field are integrated in a high-dimensional space, enabling robots to follow the desired path while simultaneously avoiding obstacles. Based on this formulation, a unified and distributed coordination and safety framework is further developed to enable cooperative motion while simultaneously incorporating obstacle avoidance and inter-robot collision prevention. Through the introduction of appropriately designed virtual coordinates, the composite guiding vector field is shown to be system-level singularity-free in the sense of collective non-vanishing and non-persistence of individual vector-field zeros. Rigorous theoretical analysis establishes safety under non-conflicting simultaneous safety constraints and proves conditional convergence when the avoidance terms eventually become inactive. Extensive numerical simulations and software-in-the-loop (SITL) experiments validate the effectiveness of the proposed method in both single-robot and multi-robot scenarios. In addition, real-world vertical takeoff and landing (VTOL) unmanned aerial vehicle (UAV) experiments empirically demonstrate the applicability of the proposed approach in realistic flight environments.
Underactuated robotic hands offer high adaptability and control simplicity, yet limited dexterity often constrains their manipulation capabilities. To address this limitation, this paper presents the G-raph hand, a reconfigurable anthropomorphic robotic hand designed for stable in-hand manipulation while maintaining control simplicity. Inspired by human manipulation, the design integrates underactuated fingers with a reconfigurable palm featuring a central compliant mechanism. This biomimetic architecture enables both active reconfiguration and passive adaptation by modulating finger-base distribution. Furthermore, a hybrid control scheme of in-hand manipulation is developed, combining position and force-feedback control with a lightweight gait planning strategy based on rapid closure property evalu ations. By integrating motion intent and system stability, the proposed framework facilitates stable grasping and significant object reorientation. Kinematic workspace analysis and extensive multi-finger manipulation experiments demonstrate that the G raph hand and its associated control framework achieve reliable, flexible, and anthropomorphic performance in complex tasks.
To track fast-moving target in cluttered environments, a series of improvements in detection, mapping, navigation, and control are introduced in previous work to make the overall system more comprehensive. However, this separated pipeline introduces significant latency and limits the agility of quadrotors. On the contrary, we follow the design principle of “less is more”, striving to simplify the process while maintaining effectiveness. In this work, we propose an end-to-end agile tracking and navigation framework for quadrotors with an elegant structure. Importantly, leveraging the multimodal nature of navigation and detection tasks, our network maintains interpretability by explicitly integrating the independent modules of the traditional pipeline, rather than a crude action regression. In detail, we adopt a set of motion primitives as anchors to cover the searching space regarding the feasible region and potential target. Then we reformulate the trajectory optimization as regression of primitive offsets considering the safety, smoothness, and other metrics. For tracking task, the trajectories are expected to approach the target and additional class scores are predicted. Subsequently, the predictions, after compensation for the estimated lumped disturbance, are transformed into thrust and attitude as control commands for swift response. We seamlessly integrate traditional planning with data-driven learning by computing the cost gradients with respect to the trajectory parameters, as in classical optimization, and directly back-propagating them to the weights of network. This eliminates the need for expert demonstration in imitation learning and provides more direct guidance than reinforcement learning. Finally, we deploy the algorithm on a compact quadrotor and conduct real-world validations in both forest and building environments to demonstrate the efficiency of the proposed method.
This paper presents a method for enhancing error performance of AC magnetic field-based navigation by integrating techniques from multiple disciplines within the electrical and electronic fields. Although the magnetic positioning system (MPS) remains robust in harsh environments that typically impair electromagnetic signals or optical measurements, its performance can be adversely affected in metal-rich settings such as the interiors of buildings. The proposed method compensates for magnetic field distortions caused by eddy current in metallic reinforcement structures within building frameworks by employing simultaneous dual-frequency magnetic field. Each of four transmitters generates the dual-frequency sine wave magnetic field using large circular coils and MOSFET H bridges. A receiver equipped with a 3-D magnetoresistive sensor measures the signals and extracts amplitude and phase information through lock-in detection techniques. The effective navigation area is extended to several meters by utilizing a free space circular current loop model. State parameters associated with both eddy current compensation and navigation are computed via a two-stage least squares algorithm, paving the way for real-time implementation. Experimental results, computed with online measurements using a handheld receiver, indicate significant reductions in position errors.
Achieving robust agile quadrotor flight under unknown external disturbances and parametric uncertainty remains a significant challenge. Inspired by robust output regulation theory, this paper proposes ${\mathcal {I} \mathcal {M}}$-NMPO, an ${\mathcal {I}}$nternal ${\mathcal {M}}$odel-based Nonlinear Model Predictive Optimization framework for robust optimal planning and control during agile flight. The core insight is the systematic decoupling of uncertain dynamics from the optimization loop via the internal model principle (IMP), enabling real-time disturbance learning and rejection while reducing reliance on high-fidelity physical models. This allows both time-optimal planning and nonlinear model predictive control (NMPC) to operate on the same disturbance-decoupled nominal dynamics. The proposed framework comprises three main components: translational and rotational nonlinear internal model compensators for real-time disturbance rejection, an NMPC optimization stabilizer for constrained agile trajectory tracking, and a receding-horizon polynomial waypoint allocation module for efficient online time-optimal reference generation. Robust constraint satisfaction, recursive feasibility and asymptotic stability are guaranteed through rigorous theoretical analysis. The effectiveness and generalizability are validated through extensive flight experiments under various disturbances, including unknown payloads, persistent fan-induced winds, and time-varying gusts, across quadrotors with different wheelbases. The framework was successfully deployed in the 2024 DJI Robomaster Intelligent MAV Championship for Planning and Control of Quadrotors, where it completed challenging racing courses at speeds up to 23.3 m/s, ranking first and finishing in less than one-third of the time taken by the second-place.
In twin-to-twin transfusion syndrome (TTTS), abnormal vascular anastomoses in the monochorionic placenta can result in unbalanced blood flow between the two fetuses. In the current practice, the laser ablation is the gold standard for surgically treating TTTS by closing abnormal anastomoses. This surgery is minimally invasive and relies on direct visualization with a fetoscope. A successful treatment requires that all anastomoses are correctly identified and ablated. While there have been attempts for mosaicking and Simultaneous Localization and Mapping (SLAM) in other surgeries, robust monocular SLAM in low textured settings remains unsolved, limiting both intraoperative guidance and robotic assistance during fetoscopy. To tackle these challenges, this work proposes a deep feature-based SLAM framework. By introducing a novel dual-modality feature tracking strategy with bundle adjustment, the proposed approach ensures real-time camera tracking and accurate dense placental surface reconstruction for surgical navigation. The proposed method is validated on the highly realistic placenta models in Unity, custom-designed placenta phantoms, as well as on an ex vivo placenta. Our method outperforms benchmark methods in low-texture and low-resolution environments. The proposed method achieves high accuracy in sparse map reconstruction. Camera pose estimation also demonstrates consistently low translational errors across all experiments. To our knowledge, this is the first SLAM framework for TTTS fetal surgery, enabling robust 3D reconstruction and tracking. Keywords: Simultaneous Localization and Mapping (SLAM), Textureless Feature Tracking, Medical Robotics, Endoscopic Fetal Surgery enabling robust 3D reconstruction and tracking.
Consistent localization of cooperative multirobot systems during navigation presents substantial challenges. This article proposes a fault-tolerant, multimodal localization framework for multirobot systems on matrix Lie groups. We introduce novel stochastic operations to perform composition, differencing, inversion, averaging, and fusion of correlated and noncorrelated estimates on Lie groups, enabling pseudopose construction for filter updates. The method integrates a combination of proprioceptive and exteroceptive measurements from inertial, velocity, and pose (pseudopose) sensors on each robot in an extended Kalman filter (EKF) framework. The prediction step is conducted on the Lie group SE2(3) & times; R-3 & times; R-3, where each robot's pose, velocity, and inertial measurement biases are propagated. The proposed framework uses body velocity, relative pose measurements from fiducial markers, and inter-robot communication to provide scalable EKF update across the network on the Lie group SE(3) & times; R-3. A fault detection module is implemented, allowing the integration of only reliable pseudopose measurements from fiducial markers. We demonstrate the effectiveness of the method through experiments with a network of wheeled mobile robots equipped with inertial measurement units, wheel odometry, and ArUco markers. The comparison results highlight the proposed method's real-time performance, superior efficiency, reliability, and scalability in multirobot localization, making it well-suited for large-scale robotic systems.
Antipodal grasping from single-view RGB-D is challenged by occlusion and partial observability, making purely analytical inference ill-posed. We present GFLA, a single-view framework that fuses learning-based perception with analytical modeling. GFLA projects antipodal contacts to the image plane, samples grasp candidates via inverse projection, and ranks them with a force-closure metric. To compensate for the information loss inherent in single-view observations, we introduce two grasping hypothesis-guided modules: (i) a Contact Projection Detection Network (CPDN) that localizes graspable regions and predicts antipodal projections on visible surfaces; and (ii) a 3D U-Net-based Scene Completion Network (SCN) that completes geometry and provides explicit collision cues. On GraspNet-1Billion, GFLA achieves its largest improvement on the novel object set (AP 35.88%, +7.59%), demonstrating superior generalization to previously unseen object categories, while also attaining a competitive overall AP of 57.84% (+1.33%). Realrobot experiments in cluttered environments, without domain adaptation or fine-tuning, achieve grasp success rates of 95.42% for single-object scenes and 90.12% for multi-object scenes, demonstrating strong practical robustness.
The robot workspace, defining the complete set of reachable end-effector positions and orientations, is critical for functional capability. However, quantitative workspace design remains an open challenge in soft robotics. This is due to complicated implicit kinematics governed by nonlinear continuum mechanics, and prohibitive computational cost of evaluating deformations across the full actuation spectrum. This article presents a computational morphogenesis framework for automatic optimization of position and orientation workspaces in multi-chamber soft pneumatic actuators. Our approach is built on three key innovations: (i) a continuum Jacobian, defined as the derivative of the end-effector's degrees of freedom (DoFs) with respect to the actuation inputs, which transforms the workspace integral from the configuration space to the actuation space, making the volume computation analytically tractable, (ii) a second-order adjoint method to derive the analytical shape derivatives of both the displacement field and this Jacobian, providing the explicit gradient of the workspace volume with respect to the robot's morphology, and (iii) a differentiable, singularity-free geometric model, complemented by an adaptive surface reconstruction algorithm, to represent and evolve the robot's free-form shape. The framework is validated by designing multi-chamber pneumatic soft actuators, achieving an 8-fold increase in 3D positional workspace volume and a 1.6-fold increase in 2D orientational workspace volume compared to the baseline Pneu-Nets actuators.
In this paper, we present a study of a novel continuum manipulator comprising a hollow, deformable tube actuated by internal cables that are constrained to remain within the structure. The hollow design simplifies fabrication, requiring only an off-the-shelf flexible tube with simple end caps, while eliminating complex backbones or predefined routing structures. This minimal hardware enables fully contained cable routing and compact operation in cluttered environments. Unlike prior works that assume fixed cable paths or external routing, the internal routing in our system is not predefined; instead, it emerges as an implicit function of manipulator deformation. We introduce a nested differentiable modeling framework based on the Geometric Variable Strain formulation, where cable routing is computed as an inner optimization problem that minimizes the cable length within the deforming body. By applying the Implicit Function Theorem, we obtain analytical derivatives of the optimal routing parameters with respect to the manipulator state, which enables computation of the analytical Jacobian of the static residual. The resulting model enables differentiable simulation, gradient-based parameter identification, and inverse kinetostatic control. We experimentally validate our framework on single- and multi-section manipulators with one and two cables, showing an average error below 5% of the manipulator body length and a maximum error of 7.54%. Finally, we demonstrate an application of the approach for a three-cable manipulator prototype in an inspection task.
Multi-robot autonomous exploration often suffers from low efficiency and high communication overhead. To address this, we propose an Unknown-Region-Guided autonomous Exploration framework (URGE), which introduces a lightweight, sub-region-based environmental representation and information sharing scheme to enable a task allocation strategy that jointly optimizes cost and task load, ultimately achieving spatially dispersed exploration. To capture sufficient spatial information while significantly reducing inter-robot communication volume, the framework incorporates a novel Regionalized Exploration Information Map (REIM) that abstracts the environment into sub-regions with different states. Based on the REIM, a task allocation strategy formulated as a joint optimization problem is proposed. It minimizes the total path cost while balancing the task load across robots, encouraging spatially dispersed exploration and fully leveraging each robot's exploration capability. Furthermore, we extend prior route planning strategies by introducing global geometric cues. This enables more globally informed exploration routes and promotes the exploration of distant, unknown regions. The proposed framework is comprehensively evaluated using five metrics through extensive comparisons with existing methods, ablation studies, and real-world experiments. Compared with the state-of-the-art (SOTA) methods, URGE improves exploration efficiency by 13.3%–45.6% and reduces redundant exploration by 22.5%–51.5%. The evaluation results also demonstrate that URGE achieves superior cooperative performance and strong practical applicability, while maintaining low communication overhead.
Collisions of rigid-link robots and rigid environments are often modeled as instantaneous events. Under this idealization, the impact forces become impulsive and the system velocities nonsmooth. In this work, we systematically analyze pre- and post-impact velocities focusing on what we refer to as the nonsmooth impact direction (NSID). We show that it is a characteristic direction of a robotic impact and largely independent of contact properties. The results are directly applicable to large classes of backdrivable robotic systems with rigid links. We address particularities of systems with nonelastic and flexible joints, unconstrained as well as constrained systems. Further, we show that the approach direction w.r.t the NSID sets the direction of the impulsive force in frictional, inelastic impacts. The comprehensive theoretical analysis of this work supported by an experimental validation may serve as a foundation for future planning and control algorithms for various robotic impact applications. These can include humanoid locomotion on a slippery surface or repetitive hammering.
Multi-robot systems must simultaneously optimize competing objectives while maintaining coordinated behavior. Existing multi-agent reinforcement learning approaches often rely on fixed or centralized coordination, which limits adaptability and violates distributed constraints. This work introduces the Coordination-Informed Multi-Objective Reinforcement Learning (CIMORL) framework, integrating a distributed weight prediction mechanism, a privileged expert training strategy, and theoretical guarantees for Pareto-optimal solutions. We present the base CIMORL method alongside two sampling-based variants, CIMORL-TS (Tree Search) and CIMORL-MPPI (MPPI), which leverage privileged global information during training to enable fully decentralized deployment. Experimental validation in cooperative and adversarial scenarios demonstrates a 21.2% hypervolume improvement and superior policy stability compared to state-of-the-art baselines. Real-world experiments with Crazyflie drones further validate the framework's robustness in resource allocation and multi-attacker multi-defend scenarios under partial observability.
Soft grippers offer gentle interaction with objects, significantly reducing the risk of damage. Their compliance in both material and structure allows them to adapt to a wide variety of object shapes and sizes. However, the deformable nature of soft grippers poses challenges for integrating precise proprioception and tactile sensing, especially when aiming for large-area, high-resolution tactile perception. In nature, the elephant's trunk exemplifies an ideal combination of compliance and tactile sensitivity, enabling it to delicately manipulate diverse objects without causing damage while also exploring its surroundings through touch. Inspired by this, we present EleTac, a soft, vision-based tactile gripper that enables safe object grasping via a pinch-like motion and delivers high-resolution, full-surface tactile feedback for integrated proprioceptive and exteroceptive sensing. Experimental results demonstrate that even with a simple control strategy, EleTac can robustly grasp and lift a variety of objects, exhibiting strong adaptability and generalization. Furthermore, the seamless integration of grasping and tactile sensing capabilities facilitates practical applications, including exploration and excavation in granular media, as well as adaptive surface following. Overall, EleTac validates the “manipulator-as-sensor” design philosophy, achieving high-quality tactile feedback without requiring additional sensing modules. For more details, please visit the project website: https://ho-lab-jaist.github.io/eletac/.
This note corrects an arithmetic step in [1, eq. (12)] and clarifies a mild structural assumption used in the CDF dominance step (9). The qualitative conclusions of [1] are unchanged; the impact is limited to a constant in the closed-form α–(c, β, γ) tradeoff of Theorem 2 and Remark 4 [see (13)–(15)].
In recent years, multimodal locomotion capabilities have enabled robots to maneuver in both terrestrial and aerial domains. However, most of these robots are designed only for locomotion, and few possess the manipulation capabilities required for practical tasks. By adding a manipulator, ground robots can perform manipulation, and some drones with robotic arms have demonstrated aerial manipulation. Nonetheless, such multirotors cannot be directly used for manipulation on the ground, and this configuration itself is unsuitable for air-ground hybrid locomotion. This is because their thruster-centralized structure makes it difficult to achieve both sufficient degrees of freedom (DoF) for manipulation and stable motion with contact and transformation. Therefore, in this work, we develop a new multilink multirotor with thrusters on each link and capable of contact with the environments. This robot can perform terrestrial rolling locomotion, aerial flight locomotion, and manipulation in multiple environments using joint actuation. First, we introduce a minimal configuration design of the proposed robot. We also describe a kinematic model and propose a design for each component based on this model. Second, we propose a real-time control method based on nonlinear optimization that considers contact and joint motion, which can be applied to various multirotors. Third, we propose motion strategies that include contact constraints specific to air-ground hybrid multilink multirotors, and analyze the limitations of manipulation capabilities based on multi-contact model. Finally, we demonstrate a variety of motions in both domains using the implemented prototype. To the best of our knowledge, this is the first demonstration of air-ground hybrid locomotion and manipulation by a multilink multirotor.