Logistics and service operations involving parcel preparation, delivery, and unpacking from a supply point to a user's home could be carried out completely by robots in the near future, taking advantage of the capabilities of the different robot morphologies for the logistics, outdoor, and domestic environments. The use of robots for parcel delivery can contribute to the goals of sustainability and reduced emissions by exploiting their different locomotion modalities (wheeled, legged, and aerial). This article reports the development and results obtained from the first robotics hackathon celebrated as part of the European Robotics and Artificial Intelligence Network involving eight robotic platforms in three domains: 1) an industrial robotic arm for parcel preparation at the supply point, 2) a Centauro robot, a dual-arm aerial manipulator, and a wheeled-legged quadruped for parcel transportation, and 3) two humanoid robots and two commercial mobile manipulators for parcel delivery and unpacking in domestic scenarios. The article describes the joint operation and the evaluation scenario, the features and capabilities of the robots, particularly those involved in the realization of the tasks, and the lessons learned.
Due to its key role in many real world applications such as autonomous driving, semantic segmentation often has to be applied on edge devices with limited computational resources. At runtime, the requirements for the segmentation system can change significantly depending on the given circumstances. For example, in some cases the focus is on a fast segmentation of certain particularly important classes, in other cases it is more relevant to distinguish a large number of classes. Besides many works aiming at a resource efficient but still accurate semantic segmentation, the possibility to adapt the segmentation model to specific circumstances is missing. In this work, we present a novel approach for flexible semantic segmentation that builds on recent developments in segmentation foundation models and prompt tuning. The approach offers the possibility to trade off inference time and accuracy, to flexibly select the classes to be segmented, and to add new target classes to an existing system just by transferring new prompts. Evaluation on an edge device shows that the inference time can be significantly reduced with fewer classes and smaller prompts, and that accuracy increases with larger prompts at the expense of a longer inference time.
This paper presents a novel approach for refining the semantic granularity of remote sensing image segmentation datasets using openly available cadastral data, thus minimizing the annotation effort. A two-step method is proposed: (1) a supervised learning step for training a semantic segmentation model on fine-grained cadastre labels to extract class prototypes, and (2) an unsupervised learning step utilizing hierarchical clustering on the extracted prototypes to generate a class hierarchy. The method is demonstrated on building segmentation in aerial imagery data, resulting in models capable of predicting new aggregated semantic classes without extra effort in dataset annotation. Evaluation of the proposed method shows its ability to generate meaningful hierarchical relationships among labels and achieve high segmentation performance.
Future crewed missions beyond low earth orbit will greatly rely on the support of robotic assistance platforms to perform inspection and manipulation of critical assets. This includes crew habitats, landing sites or assets for life support and operation. Maintenance and manipulation of a crewed site in extraterrestrial environments is a complex task and the system will have to face different challenges during operation. While most may be solved autonomously, in certain occasions human intervention will be required. The telerobotic demonstration mission, Surface Avatar, led by the German Aerospace Center (DLR), with partner European Space Agency (ESA), investigates different approaches offering astronauts on board the International Space Station (ISS) control of ground robots in representative scenarios, e.g. a Martian landing and exploration Site. In this work we present a feasibility study on how to integrate auditory information into the mentioned application. We will discuss methods for obtaining audio information and localizing audio sources in the environment, as well as fusing auditory and visual information to perform state estimation based on the gathered data. We demonstrate our work in different experiments to show the effectiveness of utilizing audio information, the results of spectral analysis of our mission assets, and how this information could help future astronauts to argue about the current mission situation.
Crewed missions to celestial bodies such as Moon and Mars are in the focus of an increasing number of space agencies. Precautions to ensure a safe landing of the crew on the extraterrestrial surface, as well as reliable infrastructure on the remote location, for bringing the crew back home are key considerations for mission planning. The European Space Agency (ESA) identified in its Terrae Novae 2030+ roadmap, that robots are needed as precursors and scouts to ensure the success of such missions. An important role these robots will play, is the support of the astronaut crew in orbit to carry out scientific work, and ultimately ensuring nominal operation of the support infrastructure for astronauts on the surface. The METERON SUPVIS Justin ISS experiments demonstrated that supervised autonomy robot command can be used for executing inspection, maintenance and installation tasks using a robotic co-worker on the planetary surface. The knowledge driven approach utilized in the experiments only reached its limits when situations arise that were not anticipated by the mission design. In deep space scenarios, the astronauts must be able to overcome these limitations. An approach towards more direct command of a robot was demonstrated in the METERON ANALOG-1 ISS experiment. In this technical demonstration, an astronaut used haptic telepresence to command a robotic avatar on the surface to execute sampling tasks. In this work, we propose a system that combines supervised autonomy and telepresence by extending the knowledge driven approach. The knowledge management is based on organizing the prior knowledge of the robot in an object-centered context. Action Templates are used to define the knowledge on the handling of the objects on a symbolic and geometric level. This robot-agnostic system can be used for supervisory command of any robotic coworker. By integrating the robot itself as an object into the object-centered domain, robot-specific skills and (tele-)operation modes can be injected into the existing knowledge management system by formulating respective Action Templates. In order to efficiently use advanced teleoperation modes, such as haptic telepresence, a variety of input devices are integrated into the proposed system. This work shows how the integration of these devices is realized in a way that is agnostic to the input devices and operation modes. The proposed system is evaluated in the Surface Avatar ISS experiment. This work shows how the system is integrated into a Robot Command Terminal featuring a 3-Degree-of-Freedom Joystick and a 7-Degree-of-Freedom haptic input device in the Columbus module of the ISS. In the preliminary experiment sessions of Surface Avatar, two astronauts on orbit took command of the humanoid service robot Rollin’ Justin in Germany. This work presents and discusses the results of these ISS-to-ground sessions and derives requirements for extending the scalable autonomy system for the use with a heterogeneous robotic team.
To become helpful assistants in our daily lives, robots must be able to understand the effects of their actions on their environment. A modern approach to this is the use of a physics simulation, where often very general simulation engines are utilized. As a result, specific modeling features, such as multi-contact simulation or fluid dynamics, may not be well represented. To improve the representativeness of simulations, we propose a framework for combining estimations of multiple heterogeneous simulations into a single one. The framework couples multiple simulations and reorganizes them based on semantically annotated action sequence information. While each object in the scene is always covered by a simulation, this simulation responsibility can be reassigned on-line. In this paper, we introduce the concept of the framework, describe the architecture, and demonstrate two example implementations. Eventually, we demonstrate how the framework can be used to simulate action executions on the humanoid robot Rollin' Justin with the goal to extract the semantic state and how this information is used to assess whether an action sequence is executed successful or not.
In future Mars exploration scenarios, astronauts orbiting the planet will control robots on the surface with supervised autonomy to construct infrastructure necessary for human habitation. Symbol-based planning enables intuitive supervised teleoperation by presenting relevant action possibilities to the astronaut. While our initial analog experiments aboard the International Space Station (ISS) proved this scenario to be very effective, the complexity of the problem puts high demands on domain models. However, the symbols used in symbolic planning are error-prone as they are often hand-crafted and lack a mapping to actual sensor information. While this may lead to biased action definitions, the lack of feedback is even more critical. To overcome these issues, this paper explores the possibility of learning the mapping between multi-modal sensor information and high-level preconditions and effects of robot actions. To achieve this, we propose to utilize a Multi-modal Latent Dirichlet Allocation (MLDA) for unsupervised symbol emergence. The learned representation is used to identify domain-specific design flaws and assist in supervised autonomy robot operation by predicting action feasibility and assessing the execution outcome. The approach is evaluated in a realistic telerobotics experiment conducted with the humanoid robot Rollin' Justin.
We present WUP-CD, a high-resolution aerial dataset for building change detection consisting of both orthorectified aerial imagery and corresponding elevation information ac-quired at two points in time. Detailed analysis of the dataset using several state-of-the-art deep learning methods allows us to show that the best results are obtained using elevation data only, highlighting its importance for future data acquisition and model development. The dataset is available for down-load 1 1 https://github.com/tritolol/WUP-CD.
Interest in accurate semantic geospatial data has increased with application areas such as the creation of realistic virtual worlds or autonomous driving. Curb detection is an important component that is especially needed for modeling road sections. By leveraging static point cloud data, we show the potential of modern 3D deep learning methods on this problem and present a post-processing step that reliably detects complex curb pathways and improves the overall result. Since datasets used in related work are mostly unavailable or not complex enough, we provide a dataset to fill this gap and serve as a basis for future work. We show that our approach generalizes well to unknown data, underlining its usefulness in real-world applications.
Certain telerobotic applications, including telerobotics in space, pose particularly demanding challenges to both technology and humans. Traditional bilateral telemanipulation approaches often cannot be used in such applications due to technical and physical limitations such as long and varying delays, packet loss, and limited bandwidth, as well as high reliability, precision, and task duration requirements. In order to close this gap, we research model-augmented haptic telemanipulation (MATM) that uses two kinds of models: a remote model that enables shared autonomous functionality of the teleoperated robot, and a local model that aims to generate assistive augmented haptic feedback for the human operator. Several technological methods that form the backbone of the MATM approach have already been successfully demonstrated in accomplished telerobotic space missions. On this basis, we have applied our approach in more recent research to applications in the fields of orbital robotics, telesurgery, caregiving, and telenavigation. In the course of this work, we have advanced specific aspects of the approach that were of particular importance for each respective application, especially shared autonomy, and haptic augmentation. This overview paper discusses the MATM approach in detail, presents the latest research results of the various technologies encompassed within this approach, provides a retrospective of DLR's telerobotic space missions, demonstrates the broad application potential of MATM based on the aforementioned use cases, and outlines lessons learned and open challenges.
Large-scale space structures, such as telescopes or spacecrafts, require suitable in-situ assembly technologies in order to overcome the limitations on payload size and mass of current launch vehicles. In many application scenarios, manual assembly by astronauts is either highly cost-inefficient or not feasible at all due to orbital constraints. However, (semi-) autonomous robotic assembly systems may provide the means to construct larger structures in space in the near future. Modularity is a key concept for such structures, and also for reducing costs in novel spacecraft designs. The advantage of the modular approach lies in the capability to generate a high number of unique assets from a reduced number of building blocks. Thus, spacecrafts can be easily adapted to particular use cases, and could even be reconfigured during their lifetime using a robotic manipulation system. These ideas lie at the core of our current EU project MOSAR (MOdular Spacecraft Assembly and Reconfiguration). Teleoperating a space robotic system from Earth to assemble a modular structure is not straightforward. Major difficulties are related to time delays, communication losses, limited control modalities, and low immersion for the operator. Autonomous robotic operations are then preferred, and with this goal we propose a fully autonomous system for planning in-space assembly tasks. Our system is able to generate assembly and reconfiguration plans for modular structures in terms of high-level actions that can autonomously be executed by a robot. Through multiple simulation layers, the system automatically verifies the feasibility and correctness of action sequences created by the planner. The layers implement different levels of abstraction, hierarchically stacked to detect infeasible transitions and initiate replanning at an early stage. Levels of abstraction increase in complexity, ranging from a basic geometric description of the spacecraft, over kinematics of the robotic setup, to full representations of the actions. The system reuses information from failed checks in all layers to avoid similar situations during replanning. We use a hybrid approach where symbolic reasoning is combined with considerations of physical constraints to generate a holistic sequence of actions. We demonstrate our planner for large space structures in a simulation environment. In particular, we consider the reconfiguration of a given modular structure, i.e. disassemble parts and reassemble them in a new configuration. The adaptability of our planning system is shown by executing the assembly plans on robots with different sets of skills and in scenarios with simulated hardware failures.
Demographic change and its various implications will offer some of the biggest challenges faced by society and our health-care systems in the coming decades. While the number of people in need of caregiving is steadily growing in most industrial nations, the number of caregivers is not keeping up with this increasing demand. Robotic assistance systems have the potential to mitigate this problem an...
Nowadays, robots are mechanically able to perform highly demanding tasks, where AI-based planning methods are used to schedule a sequence of actions that result in the desired effect. However, it is not always possible to know the exact outcome of an action in advance, as failure situations may occur at any time. To enhance failure tolerance, we propose to predict the effects of robot actions by augmenting collected experience with semantic knowledge and leveraging realistic physics simulations. That is, we consider semantic similarity of actions in order to predict outcome probabilities for previously unknown tasks. Furthermore, physical simulation is used to gather simulated experience that makes the approach robust even in extreme cases. We show how this concept is used to predict action success probabilities and how this information can be exploited throughout future planning trials. The concept is evaluated in a series of real world experiments conducted with the humanoid robot Rollin' Justin.
Intelligent robotic coworkers are considered a valuable addition in many application areas. This applies not only to terrestrial domains, but also to the exploration of our solar system. As humankind moves toward an ever increasing presence in space, infrastructure has to be constructed and maintained on distant planets such as Mars. AI-enabled robots will play a major role in this scenario. The space agencies envisage robotic co-workers to be deployed to set-up habitats, energy, and return vessels for future human scientists. By leveraging AI planning methods, this vision has already become one step closer to reality. In the METERON SUPVIS Justin experiment, the intelligent robotic coworker Rollin' Justin was controlled from Astronauts aboard the International Space Station (ISS) in order to maintain a Martian mock-up solar panel farm located on Earth to demonstrate the technology readiness of the developed methods. For this work, the system is demonstrated at AAAI 2019, controlling Rollin' Justin located in Munich, Germany from Honolulu, Hawaii.
Human teleoperation of robots and autonomous operations go hand in hand in many of todays service robots.While robot teleoperation is typically performed on low to medium levels of abstraction, automated planning has to take place on a higher abstraction level, i.e. by means of semantic reasoning.Accordingly, an abstract state of the world has to be maintained in order to enable an operator to switch seamlessly between both operational modes.We propose a novel approach that combines simulation-based geometric tracking and semantic state inference by means of so called State Inference Entities to overcome this issue.The system is demonstrated in real-world experiments conducted with the humanoid robot Rollin' Justin.
Human teleoperation of robots and autonomous operations go hand in hand in today's service robots. While robot teleoperation is typically performed on low to medium levels of abstraction, automated planning has to take place on a higher abstraction level, i.e. by means of semantic reasoning. Accordingly, an abstract state of the world has to be maintained in order to enable an operator to switch seamlessly between both operational modes. We propose a novel approach that combines simulation based geometric tracking and semantic state inference by means of so called State Inference Entities to overcome this issue. We also demonstrate how Evolutionary Strategies can be employed to refine simulation parameters. All experiments are demonstrated in real-world experiments conducted with the humanoid robot Rollin' Justin.
As robots get closer to human environments, a fundamental task for the community is to design system behaviors that foster trust. In this context, we have posed the "Green Button Challenge": every robot should have a green button that, when pressed, makes the robot explain what it is doing and why, in natural language. In this paper, we motivate why explainability is important in robotics, an why explicit knowledge representations are essential to achieving it. We highlight this with a concrete proof-of-concept implementation on our humanoid space assistant Rollin' Justin, which interprets its PDDL plans to explain what it is doing and why.