Least Square Policy Iteration (LSPI) is a model-free Reinforcement Learning algorithm capable of dealing with continuous states and actions. Chebyshev polynomials are utilized as the approximator in LSPI while Kalman Filtering handles sampled, corrupted and delayed data. Since LSPI solves optimal problems, the algorithm needs to have an exploration phase in order to avoid local minima and to cope with non-stationary cost-to-go functions. The chapter investigates how often information between neighbors in cooperative Multi-Agent Systems (MAS) needs to be exchanged in order to meet a desired performance. It suggests that stabilizing upper bounds for intervals between two consecutive broadcasting instants of each individual agent, thereby giving rise to asynchronous communication. It analyses the optimal intermittent feedback problem for MASs. The chapter explains the optimal intermittent feedback …
Giovanni Ulivi合作论文数Universita degli Studi "Roma Tre"1