Accurate force sensing and robust disturbance rejection are fundamental to stable and precise robotic contact tasks. This paper proposes a combined observer-based framework for robotic force control that compensates for spatial payload inertia and rejects disturbances within the force loop. A spatial payload momentum observer (SPMO) is developed to decouple the inertial effects of the payload from the force/torque sensor. The observer leverages spatial dynamics and provides a deterministic low-pass filtered estimate of the contact wrench, avoiding noise from numerical differentiation. Integrated into a decoupled motion-force control structure, a combined observer that accounts for force subspace dynamics and sensing filtering is designed to compensate for control disturbances. This design addresses the inherent trade-off between transient response and control stability in existing force disturbance rejection approaches. Simulations and experiments under varying payloads and disturbances validate the framework. The SPMO demonstrates superior wrench estimation accuracy compared to conventional methods, while the combined observer enables fast and robust disturbance rejection without compromising transient force control performance.
The dynamics coupling between motion and force subspaces in robotic control poses significant challenges to ensuring force control robustness, particularly under large external disturbances. While actively shaping the system inertia can eliminate this coupling, it introduces additional disturbances due to modeling uncertainties and force sensing errors. Inspired by how humans naturally adjust their elbow postures to facilitate motion force operations, we propose a quadratic programming-based nullspace optimization method that minimizes dynamics coupling for redundant torque-controlled robots. Integrated into an impedance motion force control framework, our approach minimizes an objective function defined by the Frobenius norm of the projection matrix representing inertia coupling in Cartesian space, yielding human-like postures that passively decouple task dynamics. Experimental results demonstrate that the proposed nullspace optimization significantly improves force control stability and tracking performance under conditions of high friction and external disturbances, outperforming conventional motion force control with traditional nullspace tracking approaches.
For robotic automation in intelligent manufacturing, contact-rich operations such as surface wiping, adhesive removal, and precision assembly remain challenging because small geometric deviations, surface variations, and contact uncertainties limit high product quality and system reliability. This paper proposes a visual--force--pose demonstration and phase-aware learning framework, named ForceCapture-CAFE, for learning contact-rich manufacturing skills from human physical-interaction demonstrations. A lightweight handheld device, ForceCapture, is developed to collect RGB observations, end-effector pose, gripper state, and interaction wrench during operator demonstrations. To improve the reliability of force information, the device integrates a gripper self-locking mechanism and gravity compensation to reduce hand-force interference and pose-dependent wrench bias. Based on the collected demonstrations, CAFE (Contact-Aware Force-vision Experts ) models contact-rich manipulation as a phase-structured process, where a visual expert guides pre-contact motion and a force--pose expert regulates post-contact interaction. A force-driven gate detects contact transition and fuses the expert actions under a multi-rate sensing scheme. Real-world experiments demonstrate improved task robustness and contact-stage execution. Comparative and ablation studies further verify the effectiveness of visual--force--pose demonstrations and phase-aware policy decomposition. Overall, this framework establishes a practical paradigm to transfer physical intelligence into robots, driving the advancement of highly adaptable and resilient smart manufacturing.Reproducible hardware and source code are available at https://github.com/SilkyFinish/ForceCapture-CAFE.
Real-to-Sim-to-Real technique is gaining increasing interest for robotic manipulation, as it can generate scalable data in simulation while having narrower sim-to-real gap. However, previous methods mainly focused on environment-level visual real-to-sim transfer, ignoring the transfer of interactions, which could be challenging and inefficient to obtain purely in simulation especially for contact-rich tasks. We propose ExoGS, a robot-free 4D Real-to-Sim-to-Real framework that captures both static environments and dynamic interactions in the real world and transfers them seamlessly to a simulated environment. It provides a new solution for scalable manipulation data collection and policy learning. ExoGS employs a self-designed robot-isomorphic passive exoskeleton AirExo-3 to capture kinematically consistent trajectories with millimeter-level accuracy and synchronized RGB observations during direct human demonstrations. The robot, objects, and environment are reconstructed as editable 3D Gaussian Splatting assets, enabling geometry-consistent replay and large-scale data augmentation. Additionally, a lightweight Mask Adapter injects instance-level semantics into the policy to enhance robustness under visual domain shifts. Real-world experiments demonstrate that ExoGS significantly improves data efficiency and policy generalization compared to teleoperation-based baselines. Code and hardware files have been released on https://github.com/zaixiabalala/ExoGS.
In recent years, variable camber wings (VCWs) have gained significant attention in the aviation industry due to their potential to enhance fuel efficiency, reduce noise, and improve the lift-todrag ratio. Despite extensive efforts to design VCWs, achieving both large deformations and high load-bearing capacities remains challenging. This paper introduces a novel methodology for designing morphing trailing edge based on initially curved beams (ICBs) and develops a comprehensive mathematical model for its analysis and design. We perform a compliance analysis of ICBs with varying geometry to propose a conceptual design for the trailing edge structure. The flexible structure is modeled using geometrically nonlinear Euler-Bernoulli beam theory within the Frenet framework, and its validity is confirmed through finite element analysis. The structural design is formulated as a constrained optimization problem, solved with efficient numerical methods to ensure precise deformation, load-bearing capability, and low stress levels. An optimized prototype of the morphing trailing edge has been manufactured and experimentally tested, demonstrating a camber range of +/- 25 degrees, with theoretical analysis and experimental results showing high consistency.
Scaling up robotic imitation learning for real-world applications requires efficient and scalable demonstration collection methods. While teleoperation is effective, it depends on costly and inflexible robot platforms. In-the-wild demonstrations offer a promising alternative, but existing collection devices have key limitations: handheld setups offer limited observational coverage, and whole-body systems often require fine-tuning with robot data due to domain gaps. To address these challenges, we present AirExo-2, a low-cost exoskeleton system for large-scale in-the-wild data collection, along with several adaptors that transform collected data into pseudo-robot demonstrations suitable for policy learning. We further introduce RISE-2, a generalizable imitation learning policy that fuses 3D spatial and 2D semantic perception for robust manipulations. Experiments show that RISE-2 outperforms prior state-of-the-art methods on both in-domain and generalization evaluations. Trained solely on adapted in-the-wild data produced by AirExo-2, the RISE-2 policy achieves comparable performance to the policy trained with teleoperated data, highlighting the effectiveness and potential of AirExo-2 for scalable and generalizable imitation learning.
Articulated objects are commonly found in daily life. It is essential that robots can exhibit robust perception and manipulation skills for articulated objects in real-world robotic applications. However, existing methods for articulated objects insufficiently address noise in point clouds and struggle to bridge the gap between simulation and reality, thus limiting the practical deployment in real-world scenarios. To tackle these challenges, we propose a framework towards Robust Perception and Manipulation for Articulated Objects (RPMArt), which learns to estimate the articulation parameters and manipulate the articulation part from the noisy point cloud. Our primary contribution is a Robust Articulation Network (RoArtNet) that is able to predict both joint parameters and affordable points robustly by local feature learning and point tuple voting. Moreover, we introduce an articulation-aware classification scheme to enhance its ability for sim-to-real transfer. Finally, with the estimated affordable point and articulation joint constraint, the robot can generate robust actions to manipulate articulated objects. After learning only from synthetic data, RPMArt is able to transfer zero-shot to real-world articulated objects. Experimental results confirm our approach’s effectiveness, with our framework achieving state-of-the-art performance in both noise-added simulation and real-world environments. Code, data and more results can be found on the project website at https://r-pmart.github.io.
The 3D motion field on the surface of the vision-based tactile sensor contains rich tactile information and serves as the foundation for many downstream tasks. However, achieving real-time and precise reconstruction of 3D motion field is challenging. In this study, the 3D motion field is decomposed into depth and 2D motion field, and a multi-task learning model is employed to simultaneously estimate both. The approach excels in accuracy and real-time performance (9.20 ms per frame). For depth reconstruction, the model achieves an average absolute error (MAE) of 0.062 mm. For 2D motion field tracking, an effective structured marker tracking algorithm is introduced, and a marker tracking network is constructed based on it. This network features dynamic sensing fields and a distance field auxiliary task, offering strong generalization, interpretability, and ease of training. It exhibits excellent tracking performance in various deformations, with an average error of 1.044 pixels (about 0.03 mm). Finally, the strong transferability of the multi-task model is demonstrated, with a maximum decrease in depth reconstruction accuracy of 0.018 mm.
Biologically inspired design is helpful for inspiring novel product design solutions. Currently, researchers mainly focus on biological knowledge representation and retrieval methods at a macroscopic level. At the same time, they need to pay more attention to systematic analogical reasoning method research and thus require an intelligent, biologically inspired design approach. To solve these problems, this paper presents a systematic analogical reasoning approach called Transformation-Mapping-Analogy Design (TMAD) to support biologically inspired design intelligently. Our proposed method derives its intelligence by integrating functional modeling for knowledge representation, knowledge clustering for knowledge acquisition, and design analogizing for creative concept generation. The approach also naturally reduces the workload in dealing with massive knowledge and providing more accurate, helpful knowledge. A design case is given to illustrate that the proposed algorithm can successfully achieve intelligent biologically inspired design.
While humans can use parts of their arms other than the hands for manipulations like gathering and supporting, whether robots can effectively learn and perform the same type of operations remains relatively unexplored. As these manipulations require joint-level control to regulate the complete poses of the robots, we develop AirExo, a low-cost, adaptable, and portable dual-arm exoskeleton, for teleoperation and demonstration collection. As collecting teleoperated data is expensive and time-consuming, we further leverage AirExo to collect cheap in-the-wild demonstrations at scale. Under our in-the-wild learning framework, we show that with only 3 minutes of the teleoperated demonstrations, augmented by diverse and extensive in-the-wild data collected by AirExo, robots can learn a policy that is comparable to or even better than one learned from teleoperated demonstrations lasting over 20 minutes. Experiments demonstrate that our approach enables the model to learn a more general and robust policy across the various stages of the task, enhancing the success rates in task completion even with the presence of disturbances. Project website: https://airexo.github.io/
Articulated objects like cabinets and doors are widespread in daily life. However, directly manipulating 3D articulated objects is challenging because they have diverse geometrical shapes, semantic categories, and kinetic constraints. Prior works mostly focused on recognizing and manipulating articulated objects with specific joint types. They can either estimate the joint parameters or distinguish suitable grasp poses to facilitate trajectory planning. Although these approaches have succeeded in certain types of articulated objects, they lack generalizability to unseen objects, which significantly impedes their application in broader scenarios. In this paper, we propose a novel framework of Generalizable Articulation Modeling and Manipulating for Articulated Objects (GAMMA), which learns both articulation modeling and grasp pose affordance from diverse articulated objects with different categories. In addition, GAMMA adopts adaptive manipulation to iteratively reduce the modeling errors and enhance manipulation performance. We train GAMMA with the PartNet-Mobility dataset and evaluate with comprehensive experiments in SAPIEN simulation and real-world Franka robot. Results show that GAMMA significantly outperforms SOTA articulation modeling and manipulation algorithms in unseen and cross-category articulated objects. We will open-source all codes and datasets in both simulation and real robots for reproduction in the final version. Images and videos are published on the project website at: http://sites.google.com/view/gamma-articulation
To achieve precise motion force control (MFC) when a robot interacts with the environment, selecting an appropriate control approach is crucial. A task space disturbance wrench can induce coupled accelerations between different subspaces due to the inherent coupling of task space inertia. Shaping the inertia matrix can potentially reduce or amplify this coupling effect. Inspired by impedance motion force control (IMFC), which can decouple the actuation wrench and external disturbances simultaneously, we propose partially decoupled impedance motion force control (PD-IMFC) to balance the trade-off between natural inertia conservation and motion force decoupling. The control method can be readily extended to multi-task cases, i.e., the prioritized asymmetric inertia shaping, which is the least sensitive to model uncertainties and sensing noises compared to fully or symmetric partially decoupled cases. Besides showing its relation to existing MFC approaches, the proposed approach is further extended to multiple recursively decoupled subspaces for improved force and moment tracking.
Pixel-level 2D object semantic understanding is an important topic in computer vision and could help machine deeply understand objects (e.g., functionality and affordance) in our daily life. However, most previous methods directly train on correspondences in 2D images, which is end-to-end but loses plenty of information in 3D spaces. In this paper, we propose a new method on predicting image corresponding semantics in 3D domain and then projecting them back onto 2D images to achieve pixel-level understanding. In order to obtain reliable 3D semantic labels that are absent in current image datasets, we build a large scale keypoint knowledge engine called KeypointNet, which contains 103,450 keypoints and 8,234 3D models from 16 object categories. Our method leverages the advantages in 3D vision and can explicitly reason about objects self-occlusion and visibility. We show that our method gives comparative and even superior results on standard semantic benchmarks.
3D object detection has attracted much attention thanks to the advances in sensors and deep learning methods for point clouds. Current state-of-the-art methods like VoteNet regress direct offset towards object centers and box orientations with an additional Multi-Layer-Perceptron network. Both their offset and orientation predictions are not accurate due to the fundamental difficulty in rotation classification. In the work, we disentangle the direct offset into Local Canonical Coordinates (LCC), box scales and box orientations. Only LCC and box scales are regressed, while box orientations are generated by a canonical voting scheme. Finally, an LCC-aware back-projection checking algorithm iteratively cuts out bounding boxes from the generated vote maps, with the elimination of false positives. Our model achieves state-of-the-art performance on three standard real-world benchmarks: ScanNet, SceneNN and SUN RGB-D. Our code is available on https://github.com/qq456cvb/CanonicalVoting.
Point cloud analysis without pose priors is very challenging in real applications, as the orientations of point clouds are often unknown. In this paper, we propose a brand new point-set learning framework PRIN, namely, Point-wise Rotation Invariant Network, focusing on rotation invariant feature extraction in point clouds analysis. We construct spherical signals by Density Aware Adaptive Sampling to deal with distorted point distributions in spherical space. Spherical Voxel Convolution and Point Re-sampling are proposed to extract rotation invariant features for each point. In addition, we extend PRIN to a sparse version called SPRIN, which directly operates on sparse point clouds. Both PRIN and SPRIN can be applied to tasks ranging from object classification, part segmentation, to 3D feature matching and label alignment. Results show that, on the dataset with randomly rotated point clouds, SPRIN demonstrates better performance than state-of-the-art methods without any data augmentation. We also provide thorough theoretical proof and analysis for point-wise rotation invariance achieved by our methods. The code to reproduce our results will be made publicly available.
This paper presents a general one-shot object localization algorithm called OneLoc. Current one-shot object localization or detection methods either rely on a slow exhaustive feature matching process or lack the ability to generalize to novel objects. In contrast, our proposed OneLoc algorithm efficiently finds the object center and bounding box size by a special voting scheme. To keep our method scale-invariant, only unit center offset directions and relative sizes are estimated. A novel dense equalized voting module is proposed to better locate small texture-less objects. Experiments show that the proposed method achieves state-of-the-art overall performance on two datasets: OnePose dataset and LINEMOD dataset. In addition, our method can also achieve one-shot multi-instance detection and non-rigid object localization. Code repository: https://github.com/qq456cvb/OneLoc.
Keypoint detection is an essential component for the object registration and alignment. In this work, we reckon keypoint detection as information compression, and force the model to distill out important points of an object. Based on this, we propose UKPGAN, a general self-supervised 3D keypoint detector where keypoints are detected so that they could reconstruct the original object shape. Two modules: GAN-based keypoint sparsity control and salient information distillation modules are proposed to locate those important keypoints. Extensive experiments show that our keypoints align well with human annotated keypoint labels, and can be applied to SMPL human bodies under various non-rigid deformations. Furthermore, our keypoint detector trained on clean object collections generalizes well to real-world scenarios, thus further improves geometric registration when combined with off-the-shelf point descriptors. Repeatability experiments show that our model is stable under both rigid and non-rigid transformations, with local reference frame estimation. Our code is available on https://github.com/qq456cvb/UKPGAN.
In this paper, we tackle the problem of category-level 9D pose estimation in the wild, given a single RGB-D frame. Using supervised data of real-world 9D poses is tedious and erroneous, and also fails to generalize to unseen scenarios. Besides, category-level pose estimation requires a method to be able to generalize to unseen objects at test time, which is also challenging. Drawing inspirations from traditional point pair features (PPFs), in this paper, we design a novel Category-level PPF (CPPF) voting method to achieve accurate, robust and generalizable 9D pose estimation in the wild. To obtain robust pose estimation, we sample numerous point pairs on an object, and for each pair our model predicts necessary SE(3)-invariant voting statistics on object centers, orientations and scales. A novel coarse-to-fine voting algorithm is proposed to eliminate noisy point pair samples and generate final predictions from the population. To get rid of false positives in the orientation voting process, an auxiliary binary disambiguating classification task is introduced for each sampled point pair. In order to detect objects in the wild, we carefully design our sim-to-real pipeline by training on synthetic point clouds only, unless objects have ambiguous poses in geometry. Under this cir-cumstance, color information is leveraged to disambiguate these poses. Results on standard benchmarks show that our method is on par with current state of the arts with real-world training data. Extensive experiments further show that our method is robust to noise and gives promising results under extremely challenging scenarios. Our code is available on https://github.com/qq456cvb/CPPF.
Robotic picking of diverse range of novel objects is a great challenge in dense clutter, in which objects are stacked together tightly. However, collecting large-scale dataset with dense grasp labels is extremely time-consuming, and there is a huge gap between synthetic and real images. In this paper, we explore suction based grasping from synthetic multi object rendering. To avoid tedious human labeling, we present a pipeline to model stacked objects in simulation and generate photorealistic rendering RGB-D images with dense suction point labels. To reduce simulation-to-reality gap from synthetic images to low-quality RGB-D camera, we propose a novel domain-invariant Suction Quality Neural Network (diSQNN) by training on labeled synthetic dataset with unlabeled real dataset. Specifically, diSQNN fuses photorealistic color feature and adversarial depth feature, and uses a domain discriminator on depth extractor to align depth feature distribution from synthetic and real images. We evaluate our proposed method by comparing with other baseline and suction detection method. The results demonstrate the effectiveness of our synthetic dense cluttered rendering. And through feature alignment, our domain invariant learning method can learn grasp related features, while ignoring domain related disturbing features, which maintains a high transfer performance on real RGB-D images. On a physical robot with vacuum-based gripper, the proposed method achieves average picking success rate of 91% and 88% for known objects and novel objects in a tote without using any manual labels.
Abstract This study developed somatic embryogenesis protocols for Picea pungens (Engelm), an important ornamental species, including initiation, proliferation, maturation, germination, and acclimation. Somatic embryogenic tissues were induced from mature zygotic embryos of five families, with a frequency of $$\ge $$ ≥ 22% for each. Embryogenic tissues (ET) from 13 clones of three families were proliferated for one week, achieving an average rate of 179.1%. The ET of 38 clones of three families were cultured in maturation medium for six weeks; 188 mature embryos on average were counted per gram ET cultured, of which $$\ge $$ ≥ 81.1% appeared normal, and each clone developed at least 28 normally matured embryos. A total of 69.9% or more of cotyledonary somatic embryos germinated normally and developed into normal emblings. The experiment of transplanting the emblings into a greenhouse had an average survival rate of 68.5%. Considerable variation among and within families during initiation and proliferation was observed, but this variation decreased in the maturation and germination. Changing the concentration of plant growth regulator of the initiation medium did not significantly change the initiation frequency. We recommend incorporating these protocols into the current Picea pungens practical programs, although further research is essential to increase efficiencies and reduce cost.