This book illustrates basic principles, along with the development of the advanced algorithms, to realize smart robotic systems. It speaks to strategies by which a robot (manipulators, mobile robot, quadrotor) can learn its own kinematics and dynamics from data. In this context, two major issues have been dealt with; namely, stability of the systems and experimental validations. Learning algorithms and techniques as covered in this book easily extend to other robotic systems as well. The book contains MATLAB- based examples and c-codes under robot operating systems (ROS) for experimental validation so that readers can replicate these algorithms in robotics platforms.
Sequential prediction problems such as imitation learning, where future observations depend on previous predictions (actions), violate the common i.i.d. assumptions made in statistical learning. This leads to poor performance in theory and often in practice. Some recent approaches (Daumé III et al., 2009; Ross and Bagnell, 2010) provide stronger guarantees in this setting, but remain somewhat unsatisfactory as they train either non-stationary or stochastic policies and require a large number of iterations. In this paper, we propose a new iterative algorithm, which trains a stationary deterministic policy, that can be seen as a no regret algorithm in an online learning setting. We show that any such no regret algorithm, combined with additional reduction assumptions, must find a policy with good performance under the distribution of observations it induces in such sequential settings. We demonstrate that this new approach outperforms previous approaches on two challenging imitation learning problems and a benchmark sequence labeling problem.
This paper presents a single-network adaptive critic-based controller for continuous-time systems with unknown dynamics in a policy iteration (PI) framework. It is assumed that the unknown dynamics can be estimated using the Takagi-Sugeno-Kang fuzzy model with arbitrary precision. The successful implementation of a PI scheme depends on the effective learning of critic network parameters. Network parameters must stabilize the system in each iteration in addition to approximating the critic and the cost. It is found that the critic updates according to the Hamilton-Jacobi-Bellman formulation sometimes lead to the instability of the closed-loop systems. In the proposed work, a novel critic network parameter update scheme is adopted, which not only approximates the critic at current iteration but also provides feasible solutions that keep the policy stable in the next step of training by combining a Lyapunov-based linear matrix inequalities approach with PI. The critic modeling technique presented here is the first of its kind to address this issue. Though multiple literature exists discussing the convergence of PI, however, to the best of our knowledge, there exists no literature, which focuses on the effect of critic network parameters on the convergence. Computational complexity in the proposed algorithm is reduced to the order of (Fz)(n-1) , where n is the fuzzy state dimensionality and Fz is the number of fuzzy zones in the states space. A genetic algorithm toolbox of MATLAB is used for searching stable parameters while minimizing the training error. The proposed algorithm also provides a way to solve for the initial stable control policy in the PI scheme. The algorithm is validated through real-time experiment on a commercial robotic manipulator. Results show that the algorithm successfully finds stable critic network parameters in real time for a highly nonlinear system.
This paper presents two easily implementable control schemes for balancing a cart-pole system using intelligent control tools. The proposed schemes use the Takagi–Sugeno (T–S) fuzzy model of a nonlinear system. The concept of network inversion is used to design the controller for such a system. In one of the control schemes, the control input, necessary to achieve a desired output, is computed directly through iterative inversion of the fuzzy model. In the other scheme, a parallel distributed compensator form is chosen for the controller and the parameters or the feedback gains of the controller are updated through network inversion. The updating laws are derived using both continuous and discrete time Lyapunov approaches. The inversion based parameter update avoids the need of any sufficient condition or prerequisite constraint as in existing T–S fuzzy model based control designs like LMI techniques and robust control techniques. The proposed controllers have been implemented on the cart pole system both in simulation and real time. A comparative performance analysis for different control algorithms is presented. The performance of the proposed controller is compared with the well established LQR control in real time. The proposed control algorithms work for a wide range of operating regions compared to LQR and are more robust in the sense that they can tolerate an output disturbance of higher magnitude.
We present a novel solution to the problem of robotic grasping of unknown objects using a machine learning framework and a Microsoft Kinect sensor. Using only image features, without the aid of a 3D model of the object, we implement a learning algorithm that identifies grasping regions in 2D images, and generalizes well to objects not encountered previously. Thereafter, we demonstrate the algorithm on the RGB images taken by a Kinect sensor of real life objects. We obtain the 3D world coordinates utilizing the depth sensor of the Kinect. The robot manipulator is then used to grasp the object at the grasping point.
Laxmidhar Behera合作论文数Department of Electrical Engineering17