Despite numerous improvements regarding the sample-efficiency of Reinforcement Learning (RL) methods, learning from scratch still requires millions (even dozens of millions) of interactions with the environment to converge to a high-reward policy. This is usually because the agent has no prior information about the task and its own physical embodiment. One way to address and mitigate this data-hungriness is to use Transfer Learning (TL). In this paper, we explore TL in the context of RL with the specific purpose of transferring policies from one agent to another, even in the presence of morphology discrepancies or different state-action spaces. We propose a process to leverage past knowledge from one agent (source) to speedup or even bypass the learning phase fora different agent (target) tackling the same task. Our proposed method first leverages Variational Auto-Encoders (VAE) to learn an agent-agnostic latent space from paired, time-aligned trajectories collected on a set of agents. Then, we train a policy embedded inside the created agent-invariant latent space to solve a given task, yielding a task-module reusable by any of the agents sharing this common feature space. Through several robotic tasks and heterogeneous hardware platforms, both in simulation and on physical robots, we show the benefits of our approach in terms of improved sample-efficiency. More specifically we report zero-shot generalization in some instances, where performances after transfer are recovered instantly. In worst case scenarios, performances are retrieved after fine-tuning on the target robot fora fraction of the training cost required to train a policy with similar performances from scratch.
In this paper, we present a risk-aware motion planning framework for resilient navigation in occupancy grid maps. Navigating in unknown environments involves complex interactions with the surroundings that must be effectively managed by the robot. Recently, these interactions are frequently expressed as risk constraints, where risk is defined as potential threats that could hinder the robot from accomplishing its objectives. However, when the risk constraint involves nonlinearities, common numerical solvers are severely hampered in finding a feasible solution and are prone to failure. Therefore, we present a novel risk-aware navigation strategy based on motion primitives and the Nonlinear Model Predictive Control (NMPC) method to address nonlinear risk constraints within discrete maps. We demonstrate the effectiveness of our approach through a practical application of a robust risk assessment method that takes into account both the state of the environment and the state of the robot. In addition to enhancing the decision-making capabilities of the robot, our framework offers a more resilient motion planning process that enables the robot to navigate risky scenarios where standard optimizers are likely to fail and lead to dangerous trajectories.
Despite great achievements of reinforcement learning based works, those methods are known for their poor sample efficiency. This particular drawback usually means training agents in simulated environments is the only viable option, regarding time constraints. Furthermore, reinforcement learning agents have a strong tendency to overfit on their environment, observing a drastic loss of performances at test time. As a result, tying the agent logic to its current body may very well make transfer unefficient. To tackle that issue, we propose the Universal Notice Network (UNN) method to enforce separation of the neural network layers holding information to solve the task from those related to robot properties, hence enabling easier transfer of knowledge between entities. We demonstrate the efficiency of this method on a broad panel of applications, we consider different kinds of robots, with different morphological structures performing kinematic, dynamic single and multi-robot tasks. We prove that our method produces zero shot (without additionnal learning) transfers that may produce better performances than state-of-the art approaches and show that a fast tuning enhances those performances.
Autonomous navigation in unknown 3D environments is a key issue for intelligent transportation, while still being an open problem. Conventionally, navigation risk has been focused on mitigating collisions with obstacles, neglecting the varying degrees of harm that collisions can cause. In this context, we propose a new risk-aware navigation framework, whose purpose is to directly handle interactions with the environment, including those involving minor collisions. We introduce a physically interpretable risk function that quantifies the maximum potential energy that the robot wheels absorb as a result of a collision. By considering this physical risk in navigation, our approach significantly broadens the spectrum of situations that the robot can undertake, such as speed bumps or small road curbs. Using this framework, we are able to plan safe trajectories that not only ensure safety but also actively address the risks arising from interactions with the environment.
Recent progress in robotics and artificial intelligence let us envision a future where robot presence and activity will be ubiquitous. Fueled by economics, cultural background, and design choices, human creativity will most likely design robots of various forms and shapes, using a wide range of sensors and actuators to accomplish their tasks. Consequently, it is highly probable that differently structured robots will be required to perform the same task. As many of these tasks may require learning-based control, which still relies on millions of examples to perform correctly, it would be eminently useful to be able to transfer skills from one agent to another, notwithstanding their distinct physical structure. As such, we propose a new method for the fast transfer of skills using a family of differentiable task-specific distance metrics called Task-Specific Losses (TSL). After highlighting the main shared concepts and differences with the closest existing state-of-the-art method, we demonstrate this technique on two different assistive control tasks, showing that we can indeed transfer the realization of a task learned from an expert/teacher to an agent with no previous interaction with the environment.
This letter presents a generic coordination approach for the control of mobile manipulators. It is based on an external coupled coordination and exploits a multi-trajectory following approach. A coordination stage provides optimal offsets with respect to system sub-parts reference trajectories that minimize a cost function to manage the system redundancies. In order to face the diversity of tasks to be performed by the robotic platform, the cost function is composed of different criteria that can be adapted, pending on the mission assigned to the robot. A demonstration of its generic properties is proposed on different mobile manipulator structures in simulations. Finally, the effectiveness of the proposed method is evaluated through full-scale experiments with a single-arm skid steering mobile manipulator.
General purpose simulators provide cheap training data to learn complex robotic skills. However, the transition from simulation to reality is often very challenging for the agent. One major issue is the delay on the physical robot that may deteriorate the performance of the deployed agent. Furthermore, once a successfully trained learning-based control policy is available, re-purposing the knowledge acquired by the agent to enable a structurally distinct agent to perform the same task is hazardous if done naively. In this work, we address the above issues with a single method, the DA-UNN (Delay Aware Universal Notice Network), which decomposes the knowledge into robot-specific and task-specific modules for fast transfer. Our framework deals with delays immanent to physical systems in order to improve sim2real transfer. We evaluate the efficiency of our approach using simulated and actual robots on a dynamic manipulation task where delay management is crucial.
This project addresses major problems related to the design and control of legged robots, essentially of biped type. Concerning the more specific aspect of biped robots, studies are conducted on several topics: modelling of human locomotion, in order to select the pertinent variables to be used in the control and to offer new quantitative approaches to the community of biomechanics; theoretical study of impact effects; dynamical animation; simulation analysis of a compass-like robot (chaotic behaviours have been exhibited); control of a biped robot using the task function approach. The project has also an activity on sensor-based control, especially on visual servoing.
In the last decade, robots have been taking an increasingly important place in our societies, and shall the current trend keep the same dynamic,their presence and activities will likely become ubiquitous.As robots will certainly be produced by various industrial actors, it is reasonable to assume that a very diverse robot population will be used by mankind for a broad panel of tasks.As such, it appears probable that robots with a distinct morphology will be required to perform the same task.As an important part of these tasks requires learning-based control and given the millions of interactions steps needed by these approaches to create a single agent, it appears highly desirable to be able to transfer skills from one agent to another despite a potentially different kinematic structure.Correspondingly, this paper introduces a new method, CoachGAN, based on an adversarial framework that allows fast transfer of capacities between a teacher and a student agent.The CoachGAN approach aims at embedding the teacher's way of solving the task within a critic network.Enhanced with the intermediate state variable (ISV) that translates a student state in its teacher equivalent, the critic is then able to guide the student policy in a supervised way in a fraction of the initial training time and without the student having any interaction with the target domain.To demonstrate the flexibility of this approach, CoachGAN is evaluated over a custom tennis task, using various ways to define the intermediate state variables.
Following recent trends, it appears that robot presence within human day-to-day lives is likely to grow and become ubiquitous. As many actors are engaged in this automation effort, it is plausible that the various cultural backgrounds of these actors will result in a broad range of different robots that will nevertheless need to perform similar tasks. Due to the excessively large number of experiences samples needed to successfully train a learning-based control policy, it would be remarkably useful to be able to efficiently transfer the skills acquired by a given agent to other, structurally distinct, robots. Accordingly, the BAM (Base-Abstracted Modeling) methodology proposed in this paper is a fast transfer learning approach that relies on a clear segmentation between the task model, that is a learned policy for solving a specific task and the learned robot control policy. The evaluation on two manipulation tasks using twelve different configurations of mobile manipulators demonstrates the strong potential of this approach as the segmentation results for more robust policies than naive methods and that an efficient transfer can be done in a fraction of the initial training time.
In this paper, the accurate control of mobile robots is investigated in the framework of generic edge following. It proposes a new predictive control approach, based on the minimization of the lateral error along a distance of prediction. This permits to consider a distance of convergence independently from the yaw dynamics and the longitudinal velocity, allowing to achieve harsh maneuvers, with respect to potentially kinetically unachievable paths. The proposed algorithm is generic and permits to address the control of different kinds of robots (skid-steering, car-like, etc.) in a common framework and to consider independently the speed regulation. The control proposed here allows an accurate and reactive path tracking, even if the environment is complex and narrow. The efficiency of the approach is investigated through full-scale experiments in various conditions.
In this paper, the problem associated with accurate control for mobile robots following an edge is addressed thanks to a backstepping control. In particular, the control of the angular speed (control input) is investigated through the derivation of a new backstepping control by gathering derivatives regarding to time and to the curvilinear abscissa in a single framework. This new reference then allows both time and distance convergence of the robot states towards a trajectory computed with points given by a Lidar only. This permits to address the control of different kinds of robots (skid-steering, car like, four-wheel-steering) in a common framework and to consider independently the speed regulation, the lateral and the longitudinal controls. The control proposed here allows an accurate and reactive path tracking even if the environment is complex and narrow. The efficiency of the approach is investigated through full scale experiments in various conditions.
Being able to learn and transfer skills from one agent to another is a fundamental feature in constructing even more intelligent behaviors. In this paper, we introduce a new kind of architecture and information pipeline that aims to enable the transmission of skills from one robot to one or several others. The Universal Notice Network (UNN) originality lies in the fact that it clearly distinguishes knowledge necessary to solve the task from the agent intrinsic perceptions and capabilities, hence increasing its reusability and its potential transmission to other agents. In various experiments, focusing on manipulation and comanipulation tasks in original environments, we demonstrate the capabilities of the proposed method that takes advantage of reinforcement learning algorithms and domain knowledge, such as forward geometric model and inverse kinematics. In particular, we show that a learned UNN through the interactions of an agent with its environment is transmissible to other agents, conserving a similar perfomance level.
In this paper, the problem associated with accurate control for mobile robots following an edge is addressed thanks to a backstepping control. In particular, the control of the angular speed (control input) is investigated through the derivation of a new backstepping control by gathering derivatives regarding to time and to the curvilinear abscissa in a single framework. This new reference then allows both time and distance convergence of the robot states towards a trajectory computed with points given by a Lidar only. This permits to address the control of different kinds of robots (skid-steering, car like, four-wheel-steering) in a common framework and to consider independently the speed regulation, the lateral and the longitudinal controls. The control proposed here allows an accurate and reactive path tracking even if the environment is complex and narrow. The efficiency of the approach is investigated through full scale experiments in various conditions.
In this paper, a multi-agent probabilistic optimization algorithm is applied to the problem of multi-vehicle coordination. The algorithm is known as "Probability Collectives" (PC) and has roots in Game Theory and Optimization theory. It is traditionally used for finding optimal solutions of NP-hard problems such as the travelling salesman problem. On the other end, the proposed PC formulation presented in this paper focuses on a minimal complexity implementation for solving the coordination problem in a time of the order of magnitude of 0.1. Besides time constraints, the emphasis in the design is put on ensuring that the algorithm always comes up with a feasible solution. Simulations show that both objectives are reached while having a decentralized algorithm, and flexible with respect to the type of situations it can deal with. Additional benefits of the PC algorithm include robustness to agent failure and the possibility to accommodate non-collaborative vehicles (market penetration of autonomous vehicles < 100%).
This paper proposes a path tracking strategy for wheeled mobile robots of type {1, 2} (i.e. equipped with two steering axles), with the aim to ensure the convergence of the front and rear control points along a same trajectory, leading to reduce the required space to achieve maneuvers. The proposed approach considers front and rear steering axles as two separate systems with their own control variables: the front and the rear steering angles. The problem of managing two steering axles is solved without considering an explicit control of the robot's orientation, nor a relationship between the two steering angles which is generally a not optimal approach. The proposed control laws are based on adaptive and predictive control techniques in order to address phenomena acting when moving in unstructured context, such as bad grip conditions, low-level and inertial delays. As a result, this control algorithm enables to accurately control bi-steerable mobile robots, while increasing their maneuverability. This is particularly suitable for off-road applications, such as in agriculture where potentially large robots have to move in cluttered environments and face low grip conditions.
This paper proposes a path tracking control algorithm dedicated to off-road mobile robots equipped with two steering axles. Four wheel steering mobile robot allows to enhanced motion capability with respect to classical car-like mobile robot, while reducing the friction induced by the skid-steered architecture. In particular, such a kinematic structure makes it possible to independently control the robot heading and its position, or to increase turning capabilities. In this paper, a strategy is developed in order to reduce the space required when achieving manoeuvres around a desired trajectory. Contrarily to classical point of view, expressing a relationship between the front and rear wheels, the control laws here proposed aim at following the same path for the front and the rear centre of the axle. The robot is then split into two subsystems, regulating two lateral deviations with respect to a desired trajectory.
In this paper, the problem associated with accurate control of a two-wheel steering mobile robot following a path is addressed thanks to a backstepping control strategy. This approach involves an observer to estimate the grip conditions, based on previous work, and the proposed control algorithm for the front axle. Since the significant parameters of the grip conditions are available from the observer, namely the sideslip angles and the cornering stiffnesses, it is then suitable to include them into an algorithm to control mobile robots and obtain a more accurate path tracking. This is made possible by gathering into a single backstepping approach both kinematic and dynamic models. This new point of view permits to take account of both kinematic and dynamic behaviors and grip parameters in the control law. The proposed approach is experimentally evaluated at different speeds and compared with two other state-of-the-art path tracking algorithms and evaluated for several values of lateral deviations.