Identifying predictive world models for robots from sparse online observations is essential for robot task planning and execution in novel environments. However, existing methods that leverage differentiable programming to identify world models are incapable of jointly optimizing the geometry, appearance, and physical properties of the scene. In this work, we introduce a novel rigid object representation that allows the joint identification of these properties. Our method employs a novel differentiable point-based geometry representation coupled with a grid-based appearance field, which allows differentiable object collision detection and rendering. Combined with a differentiable physical simulator, we achieve end-to-end optimization of world models or rigid objects, given the sparse visual and tactile observations of a physical motion sequence. Through a series of world model identification tasks in simulated and real environments, we show that our method can learn both simulation- and rendering-ready rigid world models from only one robot action sequence.
Camera-to-robot (also known as eye-to-hand) calibration is a critical component of vision-based robot manipulation. Traditional marker-based methods often require human intervention for system setup. Furthermore, existing autonomous markerless calibration methods typically rely on pre-trained robot tracking models that impede their application on edge devices and require fine-tuning for novel robot embodiments. To address these limitations, this paper proposes a model-based markerless camera-to-robot calibration framework, ARC-Calib, that is fully autonomous and generalizable across diverse robots and scenarios without requiring extensive data collection or learning. First, exploratory robot motions are introduced to generate easily trackable trajectory-based visual patterns in the camera's image frames. Then, a geometric optimization framework is proposed to exploit the coplanarity and collinearity constraints from the observed motions to iteratively refine the estimated calibration result. Our approach eliminates the need for extra effort in either environmental marker setup or data collection and model training, rendering it highly adaptable across a wide range of real-world autonomous systems. Extensive experiments are conducted in both simulation and the real world to validate its robustness and generalizability.
The digitization of natural history collections over the past three decades has unlocked a treasure trove of specimen imagery and metadata. There is great interest in making this data more useful by further labeling it with additional trait data, and modern "deep learning" machine learning techniques utilizing convolutional neural nets (CNNs) and similar networks show particular promise to reduce the amount of required manual labeling by human experts, making the process much faster and less expensive. However, in most cases, the accuracy of these approaches is too low for reliable utilization of the automatic labeling, typically in the range of 80-85% accuracy. In this paper, we present and validate an approach that can greatly improve this accuracy, essentially by examining the "confidence" that the network has in the generated label as well as utilizing a user-defined threshold to reject labels that fall below a chosen level. We demonstrate that a naive model that produced 86% initial accuracy can achieve improved performance - over 95% accuracy (rejecting about 40% of the labels) or over 99% accuracy (rejecting about 65%) by selecting higher confidence thresholds. This gives flexibility to adapt existing models to the statistical requirements of various types of research and has the potential to move these automatic labeling approaches from being unusably inaccurate to being an invaluable new tool. After validating the approach in a number of ways, we annotate the reproductive state of a large dataset of over 600,000 herbarium specimens. The analysis of the results points at under-investigated correlations as well as general alignment with known trends. By sharing this new dataset alongside this work, we want to allow biologists to gather insights for their own research questions, at their chosen point of accuracy/coverage trade-off.
Traditional robot control relies on analytical methods that require precise system models, which are hard to apply in real-world settings and limit generalization to arbitrary tasks. However, systems like serial manipulators and passively adaptive hands feature inherently stable regions without control discontinuities like loss of contact or singularities. In these regions, approximate controllers focusing on the correct direction of motion enable successful coarse manipulation. When coupled with a rough estimation of the motion magnitude, precision manipulation is achieved. Leveraging this insight, we introduce a novel inverse Jacobian estimation method that independently estimates the primary motion direction and magnitude of the manipulator’s actuators. Our method efficiently estimates the direct mapping from task to actuator space with no need for a priori system knowledge enabling the same framework to control both hands and arms without compromising task performance. We present a novel control method with no a priori knowledge for precision manipulation. Experiments on the Yale Model O hand, Yale Stewart Hand, and a UR5e arm demonstrate that the inverse Jacobians estimated via our approach enable real-time control with submillimeter precision in manipulation tasks. These results highlight that online self-ID data alone is sufficient for precise real-world manipulation.
Object reorientation is a key functionality in dexterous manipulation tasks, such as turning a doorknob. This is usually done on robot arms with a simple gripper and a three-degrees-of-freedom wrist. However, wrists are mechanically complex, and the wrist axes are often far away from the grasped object, resulting in coupled translations that need to be compensated with awkward whole-arm motions. We present a robot hand mechanism based on a spherical parallel architecture that can both grasp and rotate a wide range of objects in all three axes, combining much of the function of traditional wrists and grippers. The hand mechanism allows for pure spherical rotations of the grasped object about a known fixed point close to the object, thereby avoiding parasitic translations and inefficient arm motions. This point also stays fixed with respect to the hand, and is independent of the object shape, pose or initial grasp. We detail the spherical parallel design and workspace model of the wrist-like Sphinx hand, validate its performance for lower-degrees-of-freedom robot arms without traditional wrists and show that it can accurately rotate the grasped objects over large angles with basic open-loop control. Developing robot hands for unstructured human environments is a major challenge. A robotic hand that combines grasping and wrist-like rotation in one mechanism for more efficient and versatile object manipulation is presented.
Force-sensing capabilities are essential for robot manipulation systems. However, commonly used wrist-mounted force/torque sensors are heavy, fragile, and expensive, and tactile sensors require adding fragile circuitry to the robot fingers while only providing force information local to the contact. Here, we present a vision-based contact force estimator that serves as a more cost-effective and easier-to-implement alternative to existing force sensors by leveraging deformations of compliant hands upon contacts when compliant hands are in use. Our approach uses an estimator that visually observes a specialized compliant robot hand (available open source with easy fabrication through 3D printing) and predicts the contact force on the basis of its elastic deformation upon external forces. Because using wrist-mounted cameras to observe the gripper is common for robot manipulation systems, our method can obtain additional force information provided that the gripper is compliant. We optimized our compliant hand to minimize friction and avoid singularities in finger configurations, and we introduced memory to the estimator to combat the partial observability of the contact forces from the remaining friction and hysteresis. In addition, the estimator was made robust to background distractions and finger occlusions using vision foundation models to segment out the fingers. Although it is less accurate and slower than commercial force/torque sensors, we experimentally demonstrated the accuracy and robustness of our estimator (achieving between 0.2 newton and 0.4 newton error) and its utility during a variety of manipulation tasks using the gripper in the presence of noisy backgrounds and occlusions.
Joint absence in people with upper-limb-difference leads to compensatory motions. Such compensation has long been a topic of study, but typically only for a single object/user layout, which is unlikely to spatially generalize. We seek to understand how motion varies over a planar workspace for different target orientations and wrist mobility conditions. We therefore present a study that records arm and torso pose during grasping of 49 equally spaced cylindrical targets. Furthermore, we seek to validate the research practice of using wrist-immobilizing bypass sockets on able-bodied participants to simulate prostheses without wrists. Participants were 2 transradial amputees and 7 able-bodied individuals who conducted the study with and without wrist braces, generating 2450 trajectories. Heat-maps illustrate variation over the workspace in Mean Joint Angle, Range of Joint Motion and Distance Travelled by Body Segment. Results indicate that greater wrist restriction primarily exacerbated shoulder internal rotation and elbow flexion, not the trunk. We observed that bypass sockets do not fully simulate amputee behavior. Furthermore, amputee reaching with their intact limb is different to the reaching motion of normative participants, implying that transradial limb-difference affects both sides of the body. Differences in participant behavior were also observed between horizontal and vertical target orientations.
The ability to plan and control robotic in-hand manipulation is challenged by several issues, including the required amount of prior knowledge of the system and the sophisticated physics that varies across different robot hands or even grasp instances. One of the most direct models of in-hand manipulation is the inverse Jacobian, which can directly map from the desired in-hand object motions to the required hand actuator controls. However, acquiring such inverse Jacobians without complex hand-object system models is typically infeasible. We present a method for controlling in-hand manipulation using inverse Jacobians that are self-identified by a particle filter-based estimation scheme that leverages the ability of underactuated hands to maintain a passively stable grasp during self-identification movements. This method requires no a priori knowledge of the specific hand-object system and learns the system’s inverse Jacobian through small exploratory motions. Our system approximates the underlying inverse Jacobian closely, which can be used to perform manipulation tasks across a range of objects successfully. With extensive experiments on a Yale Model O hand, we show that the proposed system can provide accurate in-hand manipulation of sub-millimeter precision and that the inverse Jacobian-based controller can support real-time manipulation control of up to 900Hz.
We study the problem of rapidly identifying contact dynamics of unknown objects in partially known environments. The key innovation of our method is a novel formulation of the contact dynamics estimation problem as the joint estimation of contact geometries and physical parameters. We leverage DeepSDF, a compact and expressive neural-network-based geometry representation over a distribution of geometries, and adopt a particle filter to estimate both the geometries in contact and the physical parameters. In addition, we couple the estimator with an active exploration strategy that plans information-gathering moves to further expedite online estimation. Through simulation and physical experiments, we show that our method estimates accurate contact dynamics with fewer than 30 exploration moves for unknown objects touching partially known environments.
This systems paper presents the implementation and design of RB5, a wheeled robot for autonomous long-term exploration with fewer and cheaper sensors. Requiring just an RGB-D camera and low-power computing hardware, the system consists of an experimental platform with rocker-bogie suspension. It operates in unknown and GPS-denied environments and on indoor and outdoor terrains. The exploration consists of a methodology that extends frontier- and sampling-based exploration with a path-following vector field and a state-of-the-art SLAM algorithm. The methodology allows the robot to explore its surroundings at lower update frequencies, enabling the use of lower-performing and lower-cost hardware while still retaining good autonomous performance. The approach further consists of a methodology to interact with a remotely located human operator based on an inexpensive long-range and low-power communication technology from the internet-of-things domain (i.e., LoRa) and a customized communication protocol. The results and the feasibility analysis show the possible applications and limitations of the approach.
Precision gas analyzers are widely used in ecological research for manual measurement of soil carbon flux, a key metric used in the study of climate change. We present a generational update to the first low-cost, autonomous, closed-chamber style soil CO2 flux sensors (Fluxbots). Fluxbot 2.0 is the first such low-cost autonomous flux chamber capable of real-time wireless data transmission, which enables ecologists conducting in situ soil carbon flux surveys to set up their own wireless sensor arrays, reporting carbon flux data in real time at a very high level of temporal resolution. The system's low cost (less than 500 USD per unit) and long-range cellular data transmission capabilities also allow for greatly improved spatial resolution. Additionally, the updated system consumes significantly less power, resulting in the ability to be deployed for longer than 10x the battery lifetime of the original version on a single charge.
Due to the nature of their implementation, nearly all low-level fabrication processes produce solidly filled structures. However, lattice structures are significantly stronger for the same amount of material, resulting in structures that are much lighter and more materially efficient. Here we propose an approach for fabricating lattice structures that echoes 3D printing techniques. In it, a modular chain of specially designed links is "extruded" onto a substrate to produce various lattices configurations depending on the chosen assembly algorithm, ranging from rigid regular lattices with nodal connectivity of 12, octet-truss, to significantly less dense configurations. Compared to conventional additive manufacturing methods, our approach allows for efficient use of nearly any material or combination of materials to construct lattices with programmed arrangements. We experimentally demonstrate that a 3x3x2 lattice structure (287 total links) is fabricated in 27 minutes via a modified robotic arm and can support approximately 1000 N in compression testing.
Continuous exploration without interruption is important in scenarios such as search and rescue and precision agriculture, where consistent presence is needed to detect events over large areas. Ergodic search already derives continuous trajectories in these scenarios so that a robot spends more time in areas with high information density. However, existing literature on ergodic search does not consider the robot's energy constraints, limiting how long a robot can explore. In fact, if the robots are battery-powered, it is physically not possible to continuously explore on a single battery charge. Our paper tackles this challenge, integrating ergodic search methods with energy-aware coverage. We trade off battery usage and coverage quality, maintaining uninterrupted exploration by at least one agent. Our approach derives an abstract battery model for future state-of-charge estimation and extends canonical ergodic search to ergodic search under battery constraints. Empirical data from simulations and real-world experiments demonstrate the effectiveness of our energy-aware ergodic search, which ensures continuous exploration and guarantees spatial coverage.
Calibrating robots into their workspaces is crucial for manipulation tasks. Existing calibration techniques often rely on sensors external to the robot (cameras, laser scanners, etc.) or specialized tools. This reliance complicates the calibration process and increases the costs and time requirements. Furthermore, the associated setup and measurement procedures require significant human intervention, which makes them more challenging to operate. Using the built-in force-torque sensors, which are nowadays a default component in collaborative robots, this work proposes a self-calibration framework where robot-environmental spatial relations are automatically estimated through compliant exploratory actions by the robot itself. The self-calibration approach converges, verifies its own accuracy, and terminates upon completion, autonomously purely through interactive exploration of the environment's geometries. Extensive experiments validate the effectiveness of our self-calibration approach in accurately establishing the robot-environment spatial relationships without the need for additional sensing equipment or any human intervention.
Modular Active Cell Robots (MACROs) is a design approach in which a large number of linear actuators and passive compliant joints are assembled to create an active structure with a repeating unit cell. Such a mesh-like robotic structure can be actuated to achieve large deformation and shape-change. In this two-part paper, we use Finite Element Analysis (FEA) to model the deformation behavior of different MACRO mesh topologies and evaluate their passive and active mechanical characteristics. In part 1, we presented the passive stiffness characteristics of different MACRO meshes. Now, in this part 2 of the paper, we investigate the active strain characteristics of planar MACRO meshes. Using FEA, we quantify and compare the strains generated for the specific choice of MACRO mesh topology and further for the specific choice of actuators actuated in that particular mesh. We simulate a series of actuation modes that are based on the angular orientation of the actuators within the mesh and show that such actuation modes result in deformation that is independent of the size of the mesh. We also show that there exists a subset of such actuation modes that spans the range of deformation behavior. Finally, we compare the actuation effort required to actuate different MACRO meshes and show that the actuation effort is related to the nodal connectivity of the mesh.
Technologies from open source projects have seen widespread adoption in robotics in recent years. The rapid pace of progress in robotics is in part fueled by open source projects, providing researchers with resources, tools, and devices to implement novel ideas and approaches quickly. Open source hardware, in particular, lowers the barrier of entry to new technologies and can further accelerate innovation in robotics. But open hardware is also more difficult to propagate in comparison to open software because it involves replicating physical components, which requires users to have sufficient familiarity and access to fabrication equipment. In this work, we present a review on open robot hardware (ORH) by first highlighting the key benefits and challenges encountered by users and developers of ORH, and then relaying some best practices that can be adopted in developing successful ORH. To accomplish this, we surveyed more than 80 major ORH projects and initiatives across different domains within robotics. Finally, we identify strategies exemplified by the surveyed projects to further detail the development process, and guide developers through the design, documentation, and dissemination stages of an ORH project.
Building hand-object models for dexterous in-hand manipulation remains a crucial and open problem. Major challenges include the difficulty of obtaining the geometric and dynamical models of the hand, object, and time-varying contacts, as well as the inevitable physical and perception uncertainties. Instead of building accurate models to map between the actuation inputs and the object motions, this work proposes to enable the hand-object systems to continuously approximate their local models via a self-identification process where an underlying manipulation model is estimated through a small number of exploratory actions and non-parametric learning. With a very small number of data points, as opposed to most data-driven methods, our system self-identifies the underlying manipulation models online through exploratory actions and non-parametric learning. By integrating the self-identified hand-object model into a model predictive control framework, the proposed system closes the control loop to provide high accuracy in-hand manipulation. Furthermore, the proposed self-identification is able to adaptively trigger online updates through additional exploratory actions, as soon as the self-identified local models render large discrepancies against the observed manipulation outcomes. We implemented the proposed approach on a sensorless underactuated Yale Model O hand with a single external camera to observe the object's motion. With extensive experiments, we show that the proposed self-identification approach can enable accurate and robust dexterous manipulation without requiring an accurate system model nor a large amount of data for offline training.
Contact can be conceptualized as a set of constraints imposed on two bodies that are interacting with one another in some way. The nature of a contact, whether a point, line, or surface, dictates how these bodies are able to move with respect to one another given a force, and a set of contacts can provide either partial or full constraint on a body's motion. Decades of work have explored how to explicitly estimate the location of a contact and its dynamics, e.g., frictional properties, but investigated methods have been computationally expensive and there often exists significant uncertainty in the final calculation. This has affected further advancements in contact-rich tasks that are seemingly simple to humans, such as generalized peg-in-hole insertions. In this work, instead of explicitly estimating the individual contact dynamics between an object and its hole, we approach this problem by investigating compliance-enabled contact formations. More formally, contact formations are defined according to the constraints imposed on an object's available degrees-of-freedom. Rather than estimating individual contact positions, we abstract out this calculation to an implicit representation, allowing the robot to either acquire, maintain, or release constraints on the object during the insertion process, by monitoring forces enacted on the end effector through time. Using a compliant robot, our method is desirable in that we are able to complete industry-relevant insertion tasks of tolerances <0.25mm without prior knowledge of the exact hole location or its orientation. We showcase our method on more generalized insertion tasks, such as commercially available non-cylindrical objects and open world plug tasks.
Modular active cell robots (MACROs) are a design paradigm for modular robotic hardware that uses only two components, namely actuators and passive compliant joints. Under the MACRO approach, a large number of actuators and joints are connected to create mesh-like cellular robotic structures that can be actuated to achieve large deformation and shape change. In this two-part paper, we study the importance of different possible mesh topologies within the MACRO framework. Regular and semi-regular tilings of the plane are used as the candidate mesh topologies and simulated using finite element analysis (FEA). In Part 1, we use FEA to evaluate their passive stiffness characteristics. Using a strain-energy method, the homogenized material properties (Young's modulus, shear modulus, and Poisson's ratio) of the different mesh topologies are computed and compared. The results show that the stiffnesses increase with increasing nodal connectivity and that stretching-dominated topologies have higher stiffness compared to bending-dominated ones. We also investigate the role of relative actuator-node stiffness on the overall mesh characteristics. This analysis shows that the stiffness of stretching-dominated topologies scales directly with their cross-section area whereas bending-dominated ones do not have such a direct relationship.