This paper addresses the problem of calibrating a pair of cameras, a Microsoft Kinect sensor and an inertial measurement unit (IMU) mounted at the head of a humanoid robot with respect to its kinematic chain. As complex manipulation tasks require an accurate interplay of all involved sensors, the quality of calibration is crucial for the outcome of the intended tasks. Typical procedures for calibrating are often time-consuming, involve multiple people overseeing a series of subsequent calibration steps and require external tools. We therefore propose to auto-calibrate all sensors in a single, completely automatic and self-contained procedure, i.e. without a calibration plate. By automatically detecting a single point feature on each wrist while moving the robot’s head, the stereo cameras’, the Kinect’s infrared camera’s intrinsic and extrinsic and an IMU’s extrinsic parameters are calibrated while considering the arm joint elasticities and joint angle offsets. All parameters are obtained by formulating the calibration problem as a single least-squares batch-optimization problem. The procedure is integrated on DLR’s humanoid robot Agile Justin allowing to obtain an accurate calibration in around 5 minutes by simply “pushing a button”. The proposed approach is experimentally validated by means of standard metrics of the calibration errors.
In recent decades, the only impact of robotics on real-world applications has been confined to the execution of predetermined, repetitive tasks in controlled industrial environments. Although recent advances in all fields of robotics research have led to the development of a first generation of highly actuated, multi-sensory equipped machines, they still fall short of the range of activities humans are capable of. With the goal of having robots operate autonomously in everyday domestic environments, it is certainly necessary that human-like dynamics can be performed to a certain degree. To foster research in this direction, it is therefore often proposed to engage robots in sporting benchmark activities as these dynamic tasks are demanding for the robot’s mechanical, sensory and computational capabilities and also require a high quality of integration. This dissertation is part of the work in making a humanoid robot perform such a dynamic task, namely enabling DLR’s mobile humanoid robot Rollin’ Justin to catch up to two simultaneously thrown balls, where each ball is caught with one of its hands. To be more specific, this thesis is concerned with the perception system. Despite being a clearly defined task with easily assessable performance even for non-specialists, it is still demanding and underlines the challenges for realizing dynamic tasks in general. The challenges are: Obtain the trajectory of the thrown balls with the necessary accuracy to move the arms to the right position at the right time; handle unmodeled shaking of the robot caused by the dynamic nature of the task; avoid computational latencies while processing sensor signals to ensure proper execution within the short duration of the ball flight. From a perception point of view, this requires solving two separate problems. Firstly, for meaningful evaluation of the input data, the geometric relationships between all sensing and actuation components of the robot have to be determined through calibration. Secondly, detection, tracking, and prediction of the ball during flight have to be performed in an accurate manner while considering that the robot’s cameras also move. Of course, this has to be performed in real-time. Based on these requirements, this thesis contributes an automatic and self-contained method for calibrating all relevant sensors involved in the task. The highlights of the developed procedure are that it requires no external tools and no human assistance while achieving an accurate calibration. Furthermore, besides implementation of state-of-theart approaches for tracking balls, a general tracking scheme is proposed that integrates detection and tracking in a fully probabilistic manner. Finally, besides contributions to the task of robotic catching, this thesis further covers the work of porting the obtained methods to a ball playing entertainment robot and additional calibration problems. All presented methods and algorithms have been evaluated on the respective robots and were presented at trade fairs, public institute events and numerous lab demonstrations. Thus the methods have contributed to the development of sporting activities with humanoid robots and in doing so have extended the state of the art in service robotics.
Since its start in 1997, the setup of the RoboCup Small Size Robot League (SSL) enabled teams to use their own cameras and vision algorithms. In the fast and highly dynamic SSL environment, researchers achieved significant algorithmic advances in real-time complex colored-pattern based perception. Some teams reached, published, and shared effective solutions, but for new teams, vision processing has still been a heavy investment. In addition, it became an organizational burden to handle the multiple cameras from all the teams. Therefore, in 2008, the league started the development of a centralized, shared vision system, called SSL-Vision, which would be provided for all teams. In this paper, we discuss this system’s successful implementation in SSL itself, but also beyond it in other domains. SSL-Vision is an open source system available to any researcher interested in processing colored patterns from static cameras.
This paper presents a minimalistic robot for playing interactive ball games with human players. It is designed with a realistic entertainment application in mind, being safe, flexible, reasonably cheap, and reactive. This is achieved by a clever, minimalistic robot design with a 2 DOF roll tilt unit that moves a bat with a spherical head. The robot perceives its environment through a stereo camera system using a circle detector and a multiple hypothesis tracker. The vision system does not require a specific ball color or background structure. The paper motivates the proposed robot design with respect to the above mentioned requirements, describes our solution to the tracking, calibration, and control issues involved and presents indoor and outdoor experiments where the robot bats balls tossed by different players.
This video presents the recent upgrades of the mobile humanoid Agile Justin, bringing it closer to an ideal platform for research in autonomous manipulation. Significant upgrades have been made in the fields of mechatronics, 3D sensors, tactile skin, massive GPGPU based computing power, and software communication framework. In addition, first algorithms and two experimental scenarios are presented that take advantage of these new capabilities.
This paper addresses the problem of determining the poses of a pair of inertial sensors mounted at the opposite ends of an imperfect kinematic chain. Due to constraints during design, robots may not be equipped with the required sensors to intrinsically recover the kinematic state up to the required precision. To recover the unknown state using a pair of inertial sensors their relative pose to the opposite ends must be known. We propose an approach to calibrate these unknown relationships in a straightforward and methodically sound way while considering any existing inaccuracies in the kinematic chain. To obtain the desired parameters the calibration problem is formulated as a least squares batch-optimization problem. The proposed approach is integrated on DLR's humanoid robot Agile Justin to determine the pair of inertial sensors mounted at the opposite ends of the torso/head chain and further experimentally validated.
We believe it is possible to create the visual subsystem needed for the RoboCup 2050 challenge – a soccer match between humans and robots – within the next decade. In this position paper, we argue, that the basic techniques are available, but the main challenge will be to achieve the necessary robustness. We propose to address this challenge through the use of probabilistically modeled context, so for instance a visually indistinct circle is accepted as the ball, if it fits well with the ball’s motion model and vice versa. Our vision is accompanied by a sequence of (partially already conducted) experiments for its verification. In these experiments, a human soccer player carries a helmet with a camera and an inertial sensor and the vision system has to extract all information from that data, a humanoid robot would need to take the human’s place.
This paper presents an outdoor localization algorithm for assistive devices such as wheelchairs or walkers in urban environments. By fusing GPS, map information, and odometry with the help of a Monte Carlo particle filter, we provide adequate pose estimates for the implementation of device - specific navigation systems. We demonstrate the robustness and precision of the presented solution by experimental test runs in a municipal scenario, and compare the achieved results against a Kalman filter based localizer that integrates odometry and rate of turn data coming from a sophisticated inertial measurement unit. The core contribution of this work is given by the extension of commonly used map matching techniques in the sense that we not only evaluate the road network, but also different kinds of mapped entities representing obstacles for the vehicle.
We present a method to simultaneously track multiple objects which are subject to physical motion and can be evaluated through raw detector responses in video. Due to their two-staged design, popular tracking-by-detection approaches lack precision in the estimated trajectories due to detector inaccuracies, e.g., lighting, deformation or background clutter. Instead of separating the tasks of detection and tracking, we propose to integrate both in a single probabilistic objective function for determining the object states in a sequence. Both support each other accounting for detection inaccuracies and leading to a robust and precise single target tracker. Based on this, we extend it to multiple targets by solving the problem of determining trajectory limits and sorting out any multiple target ambiguities probabilistically. We apply our method to the task of tracking thrown balls with the goal of accurate trajectory prediction for the purpose of ball catching with a humanoid robot. Our results show improved tracking accuracy with respect to ground truth on average by around 17 %, which is dominated by increased accuracy at the beginning of the trajectory.
This paper studies different criteria for selecting configurations for the task of calibrating a robotic system. Given an automatic and self-contained procedure which allows the robot to calibrate itself without the need of external tools, we are interested in how to select the set of configurations that maximize calibration accuracy while minimizing calibration time. We experiment with the active calibration of a multi-sensorial humanoid's upper body and report that determinant-based criteria should be preferred when a greedy selection is used. In addition to criteria comparison, we further propose a new criterion for configuration selection. Its novelty stems from a direct treatment of the robot's end-effector tool variance. This is contrary to previous approaches which target the variance indirectly via calibration parameters. Our proposed objective function is derived as a compact formulation from the mean error of the robot's end-effector tool from which its variance can be computed using traditional criteria known from the theory of optimal experimental design (e.g. A-optimality).
Complex manipulation tasks require an accurate interplay of actuation and sensing. This accuracy can only be achieved by calibrating the relevant components beforehand. Typically calibration procedures are time-consuming and often include subsequent calibration steps, involve multiple people and require external tools. In this paper we alleviate these issues by auto-calibrating the different sensors of DLR's humanoid Rollin' Justin in a single, completely automatic and self-contained procedure, i.e. without calibration plate. By observing a single point feature on each wrist while moving the robot's head, the stereo cameras' intrinsic and extrinsic parameters are calibrated together with the arm joint elasticities and joint angle offsets. Additionally, we use the head motion to calibrate an Inertial Measurement Unit (IMU) extrinsically. Parameters are obtained by formulating the calibration problem as a batch-optimization problem that estimates all parameters jointly. A rough initial guess, as is, e.g., available when re-calibrating, is needed for the estimation and to facilitate marker detection. The procedure is validated on real hardware and reduces the effort considerably allowing rapid (5 min movement time), automatic, and accurate calibration by simply “pushing a button”.
The mobile humanoid Rollin'Justin is a versatile experimental platform for research in manipulation tasks. Previously, different state of the art control methods and first autonomous task execution scenarios have been demonstrated. In this video two new applications with challenging task requirements are presented. One is the catching of one or even two flying balls using all of Justin's degrees of freedom. The other is the autonomous preparation of coffee. Both applications need adequate sensors to support local referencing. The required precision in position and timing is realized in software, using the sensor information, taking the varying precision of Justin's kinematic sub-chains into account and handling all timings in sub-millisecond range.
We describe a method for estimating position and velocity of multiple flying balls for the purpose of robotic ball catching. For this a multi-target recursive Bayes filter, the Gaussian Mixture Probability Hypothesis Density filter (GMPHD), fed by a circle detector is used. This recently developed filter avoids the need to enumerate all possible data association decisions, making them computationally efficient. Over time, a mixture of Gaussians is propagated as tracks, predicted into the future and then sent to the robot. By learning a prior from training data we are focusing on detections that are likely to lead to a catchable trajectory which increases robustness. We evaluate the tracker's performance by comparing it with ground truth data, assessing tracking performance as well as the prediction precision of single tracks. Reasonable prediction performance is acquired right from the start, leading to a good overall catching rate.
This paper presents a realtime perception system for catching flying balls with DLR's humanoid Rollin' Justin. We use a two-staged bottom up approach in which we first detect balls as circles and feed these measurements into a multiple hypothesis tracker (MHT). The novel circle detection scheme works in realistic scenes without tuning parameters or background assumptions. We extend the classical multi-hypothesis tracking with prior information about the expected trajectories, therefore limiting the number of hypotheses in the first place. Since the robot starts moving while the ball is still tracked, the cameras shake heavily. A 6-DOF inertial measurements unit (IMU) is integrated to compensate this motion. Using ground-truth from a marker based tracking system we evaluate the metrical accuracy of the motion compensation as well as the tracker's prediction accuracy while in motion.
A ball catching scenario with the mobile humanoid Rollin' Justin is presented. It can catch up to two simultaneously thrown balls with its hands, reaching a catch rate of over 80%. All DOF (degrees of freedom), i.e., the arms, the torso, and the mobile platform, are used for the reaching motion and the system works completely wirelessly using only onboard sensing. The task is demanding because of the necessary precision in space (< 2cm) and time (< 5ms) as well as its realtime character due to the short flying time (< 1s). Fast perception, a good catching strategy and whole body path planning and control are important, but their tight interplay enabled by an appropriate system architecture is essential. The system overview presents the design considerations for extending Justin's versatility to this highly dynamic task. The key is not to radically change one component of the system but to do well-considered upgrades on all architectural levels, be it sensors, algorithms or middleware.
Non-linear optimization on constraint graphs has recently been applied very successfully in a variety of SLAM backends. We combine this technique with a principled way of handling non-Euclidean spaces, 3D orientations in particular, based on manifolds to build a generic and very flexible framework, the Manifold Toolkit for Matlab (MTKM). We show that MTKM makes it particularly easy to solve non-trivial multi-sensor calibration problems while remaining generic enough to handle a very different class of problems, namely SLAM, as well: After an introductory example on single camera calibration we apply MTKM to calibration of stereo vision and IMU w.r.t. the kinematic chain of a service robot, RGB-D and accelerometer calibration of a Microsoft Kinect, stereo calibration on a Nao soccer robot, and several SLAM benchmark data sets illustrating MTKM's versatility. MTKM and all presented examples are available as open source from http://openslam.org/MTK.html.
This paper shows that the field of visual SLAM has matured enough to build a visual SLAM system from open source components. The system consists of feature detection, data association, and sparse bundle adjustment. For all three modules we evaluate different libraries w.r.t. ground truth.
The current RoboCup Small Size League rules allow every team to set up their own global vision system as a primary sensor. This option, which is used by all participating teams, bears several organizational limitations and thus impairs the league’s progress. Additionally, most teams have converged on very similar solutions, and have produced only few significant research results to this global vision problem over the last years. Hence the responsible committees decided to migrate to a shared vision system (including also sharing the vision hardware) for all teams by 2010. This system – named SSL-Vision – is currently developed by volunteers from participating teams. In this paper, we describe the current state of SSL-Vision, i.e. its software architecture as well as the approaches used for image processing and camera calibration, together with the intended process for its introduction and its use beyond the scope of the Small Size League.
This paper presents a computer vision system for tracking and predicting flying balls in 3-D from a stereo-camera. It pursues a "textbook-style" approach with a robust circle detector and probabilistic models for ball motion and circle detection handled by state-of-the-art estimation algorithms. In particular we use a Multiple-Hypotheses Tracker (MHT) with an Unscented Kalman Filter (UKF) for each track, handling multiple flying balls, missing and false detections and track initiation and termination. The system also performs auto-calibration estimating physical parameters (ball radius, gravity relative to camera, air drag) simply from observing some flying balls. This reduces the setup time in a new environment.
This paper is motivated by the goal of a visual perception system for the RoboCup 2050 challenge to win against the human world-cup champion. Its contribution is to answer two questions on the subproblem of predicting the motion of a flying ball. First, if we could detect the ball in images, is that enough information to predict its motion precise enough? And second, how much do we lose by using the real-time capable Unscented Kalman Filter (UKF) instead of non-linear maximum likelihood as a gold standard? We present experiments with a camera and an inertial sensor on a helmet worn by a human soccer player. These confirm that the precision is roughly enough and using an UKF is feasible.