In this paper, we investigate a data-driven framework to solve Linear Quadratic Regulator (LQR) problems when the dynamics is unknown, with the additional challenge of providing stability certificates for the overall learning and control scheme. Specifically, in the proposed on-policy learning framework, the control input is applied to the actual (unknown) linear system while iteratively optimized. We propose a learning and control procedure, termed Relearn LQR, that combines a recursive least squares method with a direct policy search based on the gradient method. The resulting scheme is analyzed by modeling it as a feedback-interconnected nonlinear dynamical system. A Lyapunov-based approach, exploiting averaging and timescale separation theories for nonlinear systems, allows us to provide formal stability guarantees for the whole interconnected scheme. The effectiveness of the proposed strategy is corroborated by numerical simulations, where Relearn LQR is deployed on an aircraft control problem, with both static and drifting parameters.
In this paper, we consider a network of agents that jointly aim to minimise the sum of local functions subject to coupling constraints involving all local variables. To solve this problem, we propose a novel solution based on a primal-dual architecture. The algorithm is derived starting from an alternative definition of the Lagrangian function, and its convergence to the optimal solution is proved using recent advanced results in the theory of time-scale separation in nonlinear systems. The rate of convergence is shown to be linear under standard assumptions on the local cost functions. Interestingly, the algorithm is amenable to a direct implementation to deal with asynchronous communication scenarios that may be corrupted by other non-idealities such as packet loss. We numerically test the validity of our approach on a real-world application related to the provision of ancillary services in three-phase low-voltage microgrids.
In this article, we propose a distributed, first-order, feedback-based approach to solve nonlinear optimal control problems with aggregative cost functions over networks of cooperative multiagent systems. Taking inspiration from a centralized, first-order optimal control framework, named GoPRONTO, we propose a distributed method exploiting a feedback scheme iteratively updated according to a distributed tracking mechanism. Due to the aggregative structure of the problem and the desired distributed paradigm, the centralized scheme would require global quantities that are not locally available. Thus, our distributed method concurrently updates a proxy of the centralized scheme with a set of local, auxiliary variables named trackers which suitably exploit interagent communication to reconstruct the global quantities. By relying on LaSalle-based arguments, we theoretically prove that our algorithm generates a sequence of trajectories converging to the set of trajectories satisfying the first-order necessary conditions for optimality. Finally, we corroborate the theoretical results with numerical simulations on a distributed optimal control application for a fleet of 50 quadrotors.
Persistent excitation (PE) is a necessary and sufficient condition for uniform exponential parameter convergence in several adaptive, identification, and learning schemes. In this article, we consider, in the context of multi-input linear time-invariant (LTI) systems, the problem of guaranteeing PE of commonly-used regressors by applying a sufficiently rich (SR) input signal. Exploiting the analogies between time shifts and time derivatives, we state simple necessary and sufficient PE conditions for the discrete- and continuous-time frameworks. Moreover, we characterize the shape of the set of SR input signals for both single-input and multi-input systems. Finally, we show with a numerical example that the derived conditions are tight and cannot be improved without including additional knowledge of the considered LTI system.
Several interesting problems in multi-robot systems can be cast in the framework of distributed optimization. Examples include multi-robot task allocation, vehicle routing, target protection and surveillance. While the theoretical analysis of distributed optimization algorithms has received significant attention, its application to cooperative robotics has not been investigated in detail. In this paper, we show how notable scenarios in cooperative robotics can be addressed by suitable distributed optimization setups. Specifically, after a brief introduction on the widely investigated consensus optimization (most suited for data analytics) and on the partition-based setup (matching the graph structure in the optimization), we focus on two distributed settings modeling several scenarios in cooperative robotics, i.e., the so-called constraint-coupled and aggregative optimization frameworks. For each one, we consider use-case applications, and we discuss tailored distributed algorithms with their convergence properties. Then, we revise state-of-the-art toolboxes allowing for the implementation of distributed schemes on real networks of robots without central coordinators. For each use case, we discuss their implementation in these toolboxes and provide simulations and real experiments on networks of heterogeneous robots.
In this paper, we propose a novel distributed algorithm for consensus optimization over networks and a robust extension tailored to deal with asynchronous agents and packet losses. Indeed, to robustly achieve dynamic consensus on the solution estimates and the global descent direction, we embed in our algorithms a distributed implementation of the Alternating Direction Method of Multipliers (ADMM). Such a mechanism is suitably interlaced with a local proportional action steering each agent estimate to the solution of the original consensus optimization problem. First, in the case of ideal networks, by using tools from system theory, we prove the linear convergence of the scheme with strongly convex costs. Then, by exploiting the averaging theory, we extend such a first result to prove that the robust extension of our method preserves linear convergence in the case of asynchronous agents and packet losses. Further, by using the notion of Input-to-State Stability, we also guarantee the robustness of the schemes with respect to additional, generic errors affecting the agents' updates. Finally, some numerical simulations confirm our theoretical findings and compare our algorithms with other distributed schemes in terms of speed and robustness.
In this paper, we deal with a network of agents that want to cooperatively minimize the sum of local cost functions depending on a common decision variable. We consider the challenging scenario in which objective functions are unknown and agents have only access to local measurements of their local functions. We propose a novel distributed algorithm that combines a recent gradient tracking policy with an extremum-seeking technique to estimate the global descent direction. The joint use of these two techniques results in a distributed optimization scheme that provides arbitrarily accurate solution estimates through the combination of Lyapunov and averaging analysis approaches with consensus theory. We perform numerical simulations in a personalized optimization framework to corroborate the theoretical results.
This paper introduces a systematic methodological framework to design and analyze distributed algorithms for optimization and games over networks. Starting from a centralized method, we identify an aggregation function involving all the decision variables (e.g., a global cost gradient or constraint) and introduce a distributed consensus-oriented scheme to asymptotically approximate the unavailable information at each agent. Then, we delineate the proper methodology for intertwining the identified building blocks, i.e., the optimization-oriented method and the consensus-oriented one. The key intuition is to interpret the obtained interconnection as a singularly perturbed system. We rely on this interpretation to provide sufficient conditions for the building blocks to be successfully connected into a distributed scheme exhibiting the convergence guarantees of the centralized algorithm. Finally, we show the potential of our approach by developing a new distributed scheme for constraint-coupled problems with a linear convergence rate.
In this paper, we provide preliminary results toward the direction of using systems theory tools for the design and analysis of reinforcement learning algorithms. Specifically, we analyze the convergence properties of a model-based scheme with an actor-critic structure. The distinctive feature of our scheme is that the actor and critic updates are equipped with auxiliary variables that allow for the use of a constant step size. Although idealized due to the assumption on access to the underlying Markov Decision Process (MDP), the investigated setting is a starting point toward a genuine (model-free) actorcritic scheme. A key contribution is the interpretation of this algorithmic framework in terms of discrete-time, interconnected dynamical systems. Specifically, by resorting to Singular Perturbations (SP), we reinterpret the whole algorithm as the interconnection of a fast subsystem (auxiliary variables' mechanism), an intermediate one (critic), and a slow one (actor). We separately analyze three auxiliary systems each corresponding to one of the identified subsystems. These preparatory results, combined with SP and LaSalle arguments, allow us to prove that the overall method asymptotically converges to a problem stationary point. Some numerical simulations confirm our theoretical findings.
In this letter, we present a timescale separation result for discrete-time stochastic systems. We consider the feedback interconnection of two stochastic subsystems, referred to as fast and slow dynamics, and analyze them by combining timescale separation theory and stochastic LaSalle and Lyapunov theorems. Specifically, we separately focus on two auxiliary dynamics, named the boundary layer system (related to the fast part) and the reduced system (related to the slow part). For each of these auxiliary schemes, we identify a stochastic LaSalle testing condition and guarantee that satisfying both conditions is sufficient to prove almost sure LaSalle-type convergence of the original stochastic interconnection. Finally, we focus on stochastic optimization and exploit this new tool to prove almost sure convergence of the popular Stochastic Averaged Gradient and SAGA algorithms in a general nonconvex framework.
We propose fully-distributed algorithms for Nash equilibrium seeking in aggregative games over networks. We first consider the case where local constraints are present and we design an algorithm combining, for each agent, (i) the projected pseudo-gradient descent and (ii) a tracking mechanism to locally reconstruct the aggregative variable. To handle coupling constraints arising in generalized settings, we propose another distributed algorithm based on (i) a recently emerged augmented primal-dual scheme and (ii) two tracking mechanisms to reconstruct, for each agent, both the aggregative variable and the coupling constraint satisfaction. Leveraging tools from singular perturbations analysis, we prove linear convergence to the Nash equilibrium for both schemes. Finally, we run extensive numerical simulations to confirm the effectiveness of our methods and compare them with state-of-the-art distributed equilibrium-seeking algorithms.
In this paper, we propose and analyze Blockwise Incremental Gradient with Averaging (BIG-A), i.e., a novel optimization algorithm tailored for large-scale and bigdata nonconvex optimization problems with a composite cost function. At each iteration, the algorithm uses a block of a single function gradient to properly update auxiliary variables providing a proxy of a descent direction. We interpret BIGA as a dynamical system arising from the interconnection between a fast, time-varying scheme and a slow, time-invariant one. This interpretation allows us to prove the convergence properties by using system theory results relying on the LaSalleYoshizawa invariance principle and singular perturbations. The solution estimate sequence generated by BIG-A is shown to converge toward the set of stationary points of the problem, which is not assumed to be convex nor satisfying the PolyakLojasiewicz condition. If strong convexity is also assumed, linear convergence toward the unique optimal solution is established. Finally, numerical simulations confirm the theoretical findings.
We present a fully-distributed algorithm for Nash equilibrium seeking in aggregative games over networks. The proposed scheme endows each agent with a gradient-based scheme equipped with a tracking mechanism to locally reconstruct the aggregative variable, which is not available to the agents. We show that our method falls into the framework of singularly perturbed systems, as it involves the interconnection between a fast subsystem - the global information reconstruction dynamics - with a slow one concerning the optimization of the local strategies. This perspective plays a key role in analyzing the scheme with a constant stepsize, and in proving its linear convergence to the Nash equilibrium in strongly monotone games with local constraints. By exploiting the flexibility of our aggregative variable definition (not necessarily the arithmetic average of the agents' strategy), we show the efficacy of our algorithm on a realistic voltage support case study for the smart grid.
Distributed aggregative optimization is a recently emerged framework in which the agents of a network want to minimize the sum of local objective functions, each one depending on the agent decision variable (e.g., the local position of a team of robots) and an aggregation of all the agents' variables (e.g., the team barycentre). In this paper, we address a distributed feedback optimization framework in which agents implement a local (distributed) policy to reach a steady-state minimizing an aggregative cost function. We propose Aggregative Tracking Feedback, i.e., a novel distributed feedback optimization law in which each agent combines a closed-loop gradient flow with a consensus-based dynamic compensator reconstructing the missing global information. By using tools from system theory, we prove that Aggregative Tracking Feedback steers the network to a stationary point of an aggregative optimization problem with (possibly) nonconvex objective function. The effectiveness of the proposed method is validated through numerical simulations on a multi-robot surveillance scenario.
In this paper, we consider an energy trading problem in a network of interconnected MicroGrids. We consider a model in which each unit can produce, consume, or store energy and is classified as a seller or buyer, depending on its energy status. Indeed, the sellers have an excess of energy to be sold or stored, while the buyers, instead, need to buy energy from the other units or the main grid to satisfy their energy demand. In this setting, we formulate a cooperative optimization problem with the aim of finding the best tradeoff between the competitive objectives of (i) maximizing the sellers’ revenue, (ii) ensuring storage, (iii) minimizing the buyers’ energy cost, and (iv) satisfying the energy demand. Then, we recast the obtained problem in the so-called aggregative optimization scenario, a recently emerged framework in which a network of agents aims at cooperatively minimizing the sum of local functions each depending on both global (the so-called aggregative variable) and local quantities. Hence, we propose a distributed scheme tailored for aggregative optimization. The numerical simulations confirm the effectiveness of our approach showing the convergence of the chosen distributed algorithm to a stationary point of the problem. Finally, we test the flexibility of the model by considering scenarios where agents have different preferences.
In this paper, we address Linear Quadratic Regulator (LQR) problems through a novel iterative algorithm named EXtremum-seeking Policy iteration LQR (EXP-LQR). The peculiarity of EXP-LQR is that it only needs access to a truncated approximation of the infinite-horizon cost associated to a given policy. Hence, EXP-LQR does not need the direct knowledge of neither the system matrices, cost matrices, and state measurements. In particular, at each iteration, EXP-LQR refines the maintained policy using a truncated LQR cost retrieved by performing finite-time virtual or real experiments in which a perturbed version of the current policy is employed. Such a perturbation is done according to an extremum-seeking mechanism and makes the overall algorithm a time-varying nonlinear system. By using a Lyapunov-based approach exploiting averaging theory, we show that EXP-LQR exponentially converges to an arbitrarily small neighborhood of the optimal gain matrix. We corroborate the theoretical results with numerical simulations involving the control of an induction motor.
In this paper, we address multi-robot target monitoring and patrolling tasks via distributed feedback optimization. In particular, we design a distributed policy to steer a network of peer-to-peer robots toward a configuration minimizing a comprehensive index cost taking into account three different goals. First, the robots aim at disposing in a formation enclosing a given target. Second, each single robot aims at placing itself as close as possible to a specific point of interest. Third, the robots need to avoid dangerous locations modeled according to a suitable potential. To model this overall task, we resort to the formalism of aggregative feedback optimization, a recently emerged framework in which the goal is to minimize the sum of local functions each depending on both local (e.g., the position of a robot) and global variables (e.g., the barycenter of a team of robots) while concurrently taking into account also the robots’ nonlinear dynamics. We test our distributed strategy via realistic Webots simulations of a team of Crazyflie nano-quadrotors in a ROS 2 framework.
We propose a novel distributed data-driven scheme for online aggregative optimization, i.e., the framework in which agents in a network aim to cooperatively minimize the sum of local time-varying costs, each depending on a local decision variable and an aggregation of all of them. We consider a "personalized" setup in which each cost exhibits a term capturing the user's dissatisfaction and, thus, is unknown. We enhance an existing distributed optimization scheme by endowing it with a learning mechanism based on neural networks that estimate the missing part of the gradient via users' feedback about the cost. Our algorithm combines two loops with different timescales devoted to performing optimization and learning steps. In turn, the proposed scheme also embeds a distributed consensus mechanism aimed at locally reconstructing the unavailable global information due to the presence of the aggregative variable. We prove an upper bound for the dynamic regret related to (i) the initial conditions, (ii) the temporal variations of the functions, and (iii) the learning errors about the unknown cost. Finally, we test our method via numerical simulations.
In this paper, we propose a data-driven strategy to iteratively find the state feedback gain matrix solving a Linear Quadratic Regulator (LQR) problem in a model-free fashion, i.e., under unknown system and cost matrices. In our setup, we assume that, at each iteration, an oracle provides the LQR cost of the tentative policy, e.g., by running the system or a simulator. Based on this information, we develop an algorithm based on Extremum-Seeking to iteratively refine our tentative solution without any additional knowledge on the system and cost models. By using a Lyapunov-based approach exploiting averaging theory for time-varying systems, we show that the proposed algorithm exponentially converges to an arbitrarily small ball containing the optimal gain matrix. We corroborate the theoretical results by testing the proposed strategy via numerical simulations.
Sparse convex optimization involves optimization problems where the decision variables are constrained to have a certain number of entries equal to zero. In this paper, we focus on the sparse version of the so-called aggregative optimization scenario, i.e., on optimization problems in which the cost reads as the sum of local functions each depending on both a local decision variable and an aggregation of all of them. In this framework, we propose a novel fully-distributed scheme to address the problem over a network of cooperating agents. Specifically, by taking advantage of a suitable problem reformulation, we define an Augmented Lagrangian function. Then, we address such an Augmented Lagrangian by suitably interlacing the so-called Projected Aggregative Tracking distributed algorithm and the Block Coordinated Descent method giving rise to a novel fully-distributed scheme. The effectiveness of the proposed algorithm is corroborated via numerical simulations in problems arising in machine learning scenarios with both synthetic and real-world data sets.
Eduardo Camponogara合作论文数Electrical and Computer Engineering
Institute for Complex Engineered Systems2