This study introduces a decentralized control framework for a shared battery energy storage system (SBESS) serving residential households, which can select among multiple tariff schemes offered by the utility company. The problem considered is the decentralized coordination of households’ use of the SBESS, with the objective of maximizing each household’s energy arbitrage while collectively satisfying battery operating constraints and grid power exchange limits. Because households’ battery-use decisions are coupled through a shared, limited resource and future demand and photovoltaic power generation are not known a priori, this control problem is formulated within a multi-agent reinforcement learning setting, where interactions among households are modeled as a Markov game. The proposed belief-based mean-field actor–critic (BBMFAC) learning framework integrates cluster-based mean-field approximation, parameter sharing, and a semi-decentralized training and decentralized execution (s-DTDE) architecture. Within this framework, agents selecting the same tariff option are grouped into clusters and share a parameterized cluster-level critic function, thereby improving scalability and learning efficiency in large multi-agent settings. To address the inherent non-stationarity of the multi-agent environment, agents within each cluster construct beliefs from the average behaviour of other clusters during training and incorporate this information into the update of their own cluster-level critic function. The performance of the proposed method is assessed through numerical simulations using historical data on demand, photovoltaic generation, and electricity tariffs. The results are compared with state-of-the-art approaches, specifically independent and multi-agent deep deterministic policy gradient, as well as mean-field actor–critic methods.
This paper introduces a novel unsupervised method that uses advanced clustering techniques based on power and time features to identify Type 1 electrical load profiles within aggregated power measurements. The adoption of the R-statistic algorithm for the detection of ON/OFF events enhances the algorithm's performance, enabling it to capture and accurately reconstruct both slow and fast dynamic loads. A double clustering approach also guarantees that signals exhibiting identical power levels but different durations are distinctly recognised, allowing for accurate identification of individual appliances within aggregated power data. This way, the combination of clustering techniques with R-statistic improves the granularity of load profile analysis and overcomes traditional barriers in power consumption monitoring. Both simulation and experimental results are presented to evaluate and compare the performance of the proposed method to existing approaches from the literature.
This paper proposes a decentralized learning-based control scheme for solving the charging/discharging coordination problem of multiple electric vehicles (EVs) operating in a residential community, with a point of common coupling, a time-of-use tariff rate, and the uncertainty in the daily routine of the EV owners. In the proposed methodology, we model the interaction of the EV owners as a Markov Game and approach the problem through the lens of multi-agent reinforcement learning (MARL). To this end, we propose a novel belief-based Q-learning algorithm (BBQL), based on the distributed training and decentralized execution architecture. In BBQL, while training the agents form a belief about their neighbours (connected over a communication network) and simultaneously use it to better approximate their optimal action-value function. As for testing, each agent acts in a "pure" decentralized manner by playing its best-response to the developed belief. The proposed algorithm holds merit in addressing the typical key-aspects of MARL algorithms, i.e., scalability, privacy and fairness. Finally, we perform numerical simulations by using "real" historical demand, photovoltaic generation, electricity tariff and EV specification data. We compare the proposed BBQL with state-of-the-art Independent Learners. In particular, BBQL achieves a better trade-off between minimizing the electricity bill and maximizing the satisfaction of the range anxiety. Speaking statistically, BBQL results in approximate to 5 % increase in the number of days the range anxiety constraint is satisfied (with approximate to 11 % reduction in the standard deviation), while simultaneously reducing the average bill of the households by approximate to 4 % (with approximate to 13 % reduction in the standard deviation).
Sliding mode controllers (SMCs) are commonly used in permanent-magnet synchronous machines (PMSMs) for current control due to their robustness and simplicity. However, high gains used in traditional discontinuous SMC implementations can induce chattering. To address this, disturbance observers are employed to maintain robustness without resorting to high gains. This article introduces a novel continuous asymptotic SMC method for PMSM currents that avoids the need for disturbance observers, resulting in reduced complexity and tuning efforts. The control laws of the two dq-axes currents are obtained through the sensitivity of the tracking errors with respect to the controller outputs. The robustness and convergence properties of the proposed control laws are theoretically studied using the Lyapunov approach. Numerical simulations are used to evaluate the performance and robustness of the proposed controller, followed by experiments to compare it to a discontinuous terminal SMC with and without a disturbance observer. The results clearly demonstrate the superiority of the proposed controller that ensures fast convergence, low chattering, and high robustness to parameter variations without requiring the design of additional disturbance observers.
Modern wind energy conversion systems equipped with permanent magnet synchronous generators (PMSG) require continuous adaptation of rotor blade rotational speed to the wind speed to maximize power generation. However, precise rotational speed regulation is challenged by abrupt wind speed variations, along with uncertainties and changes in the characteristic electromechanical parameters. This study presents an innovative sliding mode control technique based on a continuous reaching law to ensure a precise speed regulation of PMSGs avoiding typical chattering issues of more conventional discontinuous sliding mode techniques. The stability of the proposed controller is analytically demonstrated according to the Lyapunov stability theory. Simulation and experimental results are provided to illustrate the robustness of the proposed controller and its superiority to a more conventional discontinuous sliding mode controller.
This article proposes a data-driven decentralized control scheme fora battery energy storage system, shared among residential PV households characterized by their respective uncontrollable demand and PV generation. The households are connected to the grid via the point of common coupling and are accordingly billed by the utility company. We firstly translate the decentralized control objective into a multi-agent reinforcement learning (MARL) problem by modelling the interaction between the agents and their environment as a Markov Game. Thereafter, we present the novel Distributed Subgradient Q-learners (DSQL) algorithm based on the localization of the Hyper-Q function and the coordination among the learning agents connected via a communication network. The proposed algorithm holds merit in addressing the typical key-aspects of MARL algorithms, i.e., scalability, privacy and fairness. Finally, we perform numerical simulations by using real historical demand, PV generation and electricity tariff data and highlight the key advantages of the proposed algorithm w.r.t. the state-of-art, in terms of economic savings and key-performance indicators, such as peak-to-average ratio, valley-to-average ratio and root-mean-squared-deviation.
This paper presents a data-driven practical stabilization approach for solving stochastic Dynamic Programming problems with unknown Markov Decision Process models over an infinite time horizon. The Bellman operator is modeled as a discrete-time switched affine system, with each mode representing a specific stationary stochastic policy and an external bounded disturbance term to account for such modeling issue. A two-step approach is followed. First, a model-based robust practical stabilization problem is solved to derive stabilization conditions which enable the practical convergence of the resulting closed-loop system trajectories towards a chosen reference value function. Then, by exploiting recent model-to-data Linear Matrix Inequality transformation tools, these results are further developed to obtain data-driven robust stabilization conditions for addressing the case of model-free problems. Such data-driven stabilization conditions are deployed into the Value Iteration algorithm, and finally tested on the recycling robot and the parking lot management problems to demonstrate the effectiveness of the proposed method.
This paper presents a switching control strategy as a criterion for policy selection in stochastic Dynamic Programming problems over an infinite time horizon. In particular, the Bellman operator, applied iteratively to solve such problems, is generalized to the case of stochastic policies, and formulated as a discrete-time switched affine system. Then, a Lyapunov-based policy selection strategy is designed to ensure the practical convergence of the resulting closed-loop system trajectories towards an appropriately chosen reference value function. This way, it is possible to verify how the chosen reference value function can be approached by using a stabilizing switching signal, the latter defined on a given finite set of stationary stochastic policies. Finally, the presented method is applied to the Value Iteration algorithm, and an illustrative example of a recycling robot is provided to demonstrate its effectiveness in terms of convergence performance. (c) 2024 Elsevier Ltd. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
Effective energy management in Low Earth Orbit satellites is critical, as inefficient energy management can significantly affect mission objectives. The dynamic and harsh space environment further complicates the development of effective energy management strategies. To address these challenges, we propose a Deep Reinforcement Learning approach using Deep-Q Network to develop an adaptive energy management framework for Low Earth Orbit satellites. Compared to traditional techniques, the proposed solution autonomously learns from environmental interaction, offering robustness to uncertainty and online adaptability. It adjusts to changing conditions without manual retraining, making it well-suited for handling modeling uncertainties and non-stationary dynamics typical of space operations. Training is conducted using a realistic satellite electric power system model with accurate component parameters and single-orbit power profiles derived from real space missions. Numerical simulations validate the controller performance across diverse scenarios, including multi-orbit settings, demonstrating superior adaptability and efficiency compared to conventional Maximum Power Point Tracking methods.
Dynamic Programming suffers from the curse of dimensionality due to large state and action spaces, a challenge further compounded by uncertainties in the environment. To mitigate these issue, we explore an off-policy based Temporal Difference Approximate Dynamic Programming approach that preserves contraction mapping when projecting the problem into a subspace of selected features, accounting for the probability distribution of the perturbed transition probability matrix. We further demonstrate how this Approximate Dynamic Programming approach can be implemented as a particular variant of the Temporal Difference learning algorithm, adapted for handling perturbations. To validate our theoretical findings, we provide a numerical example using a Markov Decision Process corresponding to a resource allocation problem.
Aviation significantly contributes to CO2 and NOX emissions, which, particularly at cruise altitudes, also lead to ozone production. This has driven the push for greener aviation policies. Hybrid-electric propulsion systems, which combine traditional combustion engines with electric powertrains, offer reduced fuel consumption, lower emissions, and increased range. Hydrogen fuel cells, with their high energy density, are a promising alternative to conventional batteries. However, these systems add complexity in design and management, necessitating advanced Energy Management Systems (EMSs) for optimal performance. While EMS strategies for battery-powered hybrids are well-studied, research on fuel cell-based hybrids is still limited. This paper proposes an EMS strategy using Dynamic Programming (DP) for a hybrid aircraft with a gasoline engine and a fuel cell-powered electric motor. The DP approach aims to optimize the power split between these propulsion sources to minimize fuel consumption. Simulation results across various flight scenarios demonstrate the effectiveness of this approach.
Dynamic Programming suffers from the curse of dimensionality due to large state and action spaces, a challenge further compounded by uncertainties in the environment. To mitigate these issue, we explore an off-policy based Temporal Difference Approximate Dynamic Programming approach that preserves contraction mapping when projecting the problem into a subspace of selected features, accounting for the probability distribution of the perturbed transition probability matrix. We further demonstrate how this Approximate Dynamic Programming approach can be implemented as a particular variant of the Temporal Difference learning algorithm, adapted for handling perturbations. To validate our theoretical findings, we provide a numerical example using a Markov Decision Process corresponding to a resource allocation problem.
Distributed control architectures are attractive for large-scale interconnected systems as they provide good trade-offs between control complexity and closed-loop performance. In such a context, it becomes crucial to ensure robustness against variations in the communication topology arising from connectivity failures or cyber-attacks. Based on Linear Matrix Inequalities, this paper introduces a novel design approach for distributed controllers that exhibit robustness to changes in the communication topology. The primary objective is to achieve resilience against potential disconnections of subsystems that may occur within the communication network. The control gains are structured, and reflect the nominal communication topology. Exponential stability with a prescribed decay rate is guaranteed within the sub-configurations of the nominal communication topology. The proposed approach does not make any assumptions regarding the connectivity of the communication graph. Moreover, leveraging on the projection lemma, the design of control gains is decoupled from the design of the Lyapunov matrix, thus minimizing the conservativeness of the solution. An illustrative example on a multimachine power system is presented to demonstrate the effectiveness of the proposed approach.
This article proposes a novel kernel-based Dynamic Programming (DP) approximation method to tackle the typical curse of dimensionality of stochastic DP problems over the finite time horizon. Such a method utilizes kernel functions in combination with Support Vector Machine (SVM) regression to determine an approximate cost function for the entire state space of the underlying Markov Decision Process (MDP), by leveraging cost function computed for selected representative states. Kernel functions are used to define the so-called kernel matrix, while the parameter vector of the given kernel-based cost function approximation is computed by moving backwards in time from the terminal condition and by applying SVM regression. This way, the difficulty of selecting a proper set of features is also tackled. The proposed method is then extended to the infinite time horizon case. To show the effectiveness of the proposed approach, the resulting Recursive Residual Approximate Dynamic Programming (RR-ADP) algorithm is applied to the sensor scheduling design in multi-process remote state estimation problems.
Cell balancing is used in battery systems to guarantee uniform charge and discharge of their cells during operations, and aims at improving the performance of the whole battery pack. Onboard battery performance and lifespan are particularly important in Electric Vehicles (EVs), since they have a direct impact on their autonomy. This paper proposes a Deep Reinforcement Learning (DRL)-based framework for Dynamic Reconfigurable Batteries (DRBs), where the capability of dynamically reconfiguring their cell topology can be exploited to attain cell balancing in EV applications. Thanks to the model-free nature and the robustness/adaptability properties of DRL-based solutions, the resulting trained agent is able to reach DRB cell balancing, while also taking into account both operational and modelling aspects of the combined EV-DRB system (e.g., all the events included in a typical driving cycle, the EV regenerative braking, and the heterogeneity of the DRB cells). The training process is carried out by using standard driving cycles, while the resulting trained agent is validated with a dataset based on real driving profiles. Finally, the performance achieved by the DRL policies is compared with some heuristic/rule-based approaches inspired from the literature.
This paper addresses the limited adaptability and the computational burden of energy management systems (EMSs) for hybrid electric vehicles (HEVs) implemented via dynamic programming (DP)-based approaches. First, a deterministic dynamic programming (DDP) framework is presented to solve HEV EMS problems subject to a specific driving cycle. To address this limitation, an improved DDP approach, integrating the actual travelled position of the vehicle into the control law, is proposed. This way, a given DDP-based EMS can be applied to all the driving cycles, yet still measured on the same road. Stochastic dynamic programming (SDP)-based EMSs are also developed and prove to be more adaptive to driving scenarios completely different from the ones used for their computation. Real-world driving cycles are employed in all the presented cases, while a reduced HEV powertrain model is used to alleviate the typical DP computational burden.
In this paper, we analyse the convergence properties of the Dynamic Programming Value Iteration algorithm by exploiting the stability theory of discrete-time switched affine systems. More specifically, by formulating the Value Iteration algorithm as a switched affine system, a Lyapunov-based optimal policy selection strategy is designed to guarantee the practical stabilisation of the resulting system towards an invariant set of attraction containing a given target value function. The switching control algorithm, referred to as Lyapunov-based Value Iteration algorithm, can be regarded as a convergence analysis tool and can be adopted to verify if and how such target value function can be approached by choosing from a subset of suitable stationary policies, at each time slot. The usage of the proposed algorithm in practice is also discussed. Finally, two different applications are provided to further illustrate and examine the key-aspects of the approach presented.
This paper addresses both the modeling and the resolution of the replacement problem for a population of machines. The main objective is the computation of a minimum cost replacement policy, which, based on the status of each machine, determines whether one or more machines have to be replaced over a given finite time horizon.The replacement problem of a set of machines can be regarded as a sequential decision-making problem under uncertainty. Thanks to this, we propose a novel formulation for such problems consisting of a composition of discrete-time multi-state Markov Decision Processes (MDPs), one for each specific machine. The underlying optimization problem is formulated as a stochastic Dynamic Programming (DP), and then solved by using the principles of the backward DP algorithm. Moreover, to deal with the curse of dimensionality due to the high-cardinality state-space of real-world/industrial applications, a new generalized multi-trajectory Least-Squares Temporal Difference (LSTD) based method is introduced. The resulting algorithm computes an approximate optimal cost function by: (i) running Monte Carlo simulations over different trajectories of a given length; (ii) embedding the policy improvement step within the recursive LSTD iterations; (iii) enforcing an off-policy mechanism to improve the LSTD exploration capabilities. A study on the convergence properties of the proposed approach is also provided. Several numerical examples are given to illustrate its effectiveness in terms of parametric sensitivity, computational burden, and performance of the computed policies compared with some heuristics defined in the literature.
In this letter, we study the say decentralized scheduling of an energy storage system say shared among residential households. In particular, we consider the households as learning agents and model their interaction as a Markov Game. To address the challenges associated with the non-stationary nature of the multi-agent learning, we propose a consensus-based Tabular $Q-$ learning method. Additionally, we provide simulation studies utilizing a real-world household dataset and demonstrate the effectiveness of our approach.
Power buffers are electronic converters that shield direct current (dc) microgrids from the effects of abrupt load changes. By promoting collective behaviors, distributed control solutions can improve the overall grid performance, as power buffers can collectively react to abrupt load changes. In this article, a novel gain-scheduled structured control scheme for power buffers is proposed. The control objective is to regulate both the stored energy and the input impedance of the power buffers in a distributed fashion, while guaranteeing, at the same time, the closed loop stability regardless of load changes occurring within predefined ranges. The proposed controller is expressed as a gain-scheduled state feedback law, whose coefficients depend affinely on the scheduling parameters. The structure of the controller gain is chosen to reflect the topology of the communication network implementing the distributed controller. In this article, the existing theory is extended to deal with such control problem, providing a novel strategy, which can be used to design gain-scheduled controllers for the considered class of structured feedback systems. Simulation results using a high-fidelity dc microgrid simulator verify the effectiveness of the proposed approach.