Safe reinforcement learning (SRL) aims to optimize control policies that maximize long-term reward, while adhering to safety constraints. SRL has many real-world applications such as, autonomous vehicles, industrial robotics, and healthcare. Recent advances in offline reinforcement learning (RL) - where agents learn policies from static datasets without interacting with the environment - have made it a promising approach to derive safe control policies. However, offline RL faces significant challenges, such as covariate shift and outliers in the data, which can lead to suboptimal policies. Similarly, online SRL, which derives safe policies through real-time environment interaction, struggles with outliers and often relies on unrealistic regularity assumptions, limiting its practicality. This paper addresses these challenges by proposing a hybrid-offline-online approach. First, prior knowledge from offline learning guides online exploration. Then, during online learning, we replace the popular Gaussian Process (GP) with the Student-t's Process (TP) to enhance robustness to covariate shift and outliers.
Achieving average consensus without disclosing the initial agents' state is critical for secure multi- agent coordination. This paper proposes a novel privacy-preserving average consensus algorithm via a matrix-weighted inter-agent coupling mechanism. Specifically, the algorithm first lifts each agent state to a higher-dimensional space, then employs a dedicatedly designed matrix-valued state coupling mechanism to conceal the initial agents' state while guaranteeing that the multi-agent network achieves average consensus. The convergence analysis is transformed into the average consensus problem on matrix-weighted switching networks with low-rank, positive semi-definite coupling matrices. We show that the average consensus can be guaranteed and discuss its performance in the presence of honest-but-curious agents and external eavesdroppers. The algorithm, involving only basic matrix operations, is computationally more efficient than cryptography-based approaches and can be implemented without relying on a centralized third party. Numerical results are provided to illustrate the effectiveness of the algorithm. (c) 2024 Published by Elsevier Ltd.
In this article, we consider a class of nonlinear control systems subject to false data injection attacks and switching attacks. The problem of attack estimation is formulated as the simultaneous reconstruction of system states, attack vectors, and system modes of a switched nonlinear system. In the proposed attack estimation algorithm, the inverse system of each mode aims to estimate system states and attack vectors when the corresponding mode has a finite relative order. A set-valued mode index inversion map correctly recovers the true mode when the switch-singular pair between any two modes does not exist. Two case studies, including a consensus problem and a supervised learning problem, are used to validate the performance of the developed algorithm.
A privacy-preserving average consensus algorithm is proposed that synergizes the Beaver triple in secret sharing theory and noise obfuscation. The algorithm safeguards the initial values of agents against passive adversaries in a multiagent system. It is proved that the proposed algorithm can concurrently ensure average consensus and privacy, while also reducing the online computation and communication overhead compared to encryption-based ones. In addition, it imposes a less stringent condition for privacy preservation compared to certain noise-obfuscation techniques.
This paper considers perception-driven control of a mobile robot for reference tracking where perception is performed by a machine learning system. The robot is subject to passive attacks and evasion attacks on image transmission. A robust output feedback controller together with a chaotic encryption system ensures input-to-state stability of the closed-loop system, and the chaotic encryption approach keeps image transmission secure. Simulations are conducted in the CARLA simulator to demonstrate robust reference tracking and secure image transmission.
A privacy-preserving average consensus algorithm is designed based on the Beaver triple technique against passive adversaries. The Beaver triple technique is integrated into a restructure of the discrete-time average consensus algorithm to preserve the privacy of initial values of agents in a multiagent system. The performance of the algorithm is theoretically analyzed.
In this paper, the privacy and security issues associated with the transactive energy system (TES) deployment over insecure communication links are addressed. In particular, it is ensured that (1) individual agents' bidding information is kept private throughout hierarchical market-based interactions; and (2) any extraneous data injection attack can be quickly and easily detected. An implementation framework is proposed to enable the cryptography-based enhancement of privacy and security for the deployment of any general hierarchical systems including TESs. Under the proposed framework, a unified cryptography-based approach is developed to achieve both privacy and security simultaneously. Specifically, privacy preservation is realized by an enhanced Paillier encryption scheme, where a block design is proposed to significantly improve computational efficiency. Attack detection is further achieved by an enhanced Paillier digital signature scheme, where a stamp-concatenation mechanism is proposed to enable detection of data replace and reorder attacks. Simulation results verify the effectiveness of the proposed cyber-resilient design for transactive energy systems.
This paper examines the distributed neighbor selection problem for second-order semi-autonomous multi-agent networks. By inheriting the leader-to-follower reachability property encoded in the eigenvector associated with the smallest eigenvalue of perturbed graph Laplacian, this paper shows that the convergence rate of a second-order semi-autonomous network can be enhanced on the reduced network constructed by this eigenvector. Moreover, a quantitative connection between the relative rate of change in velocity of neighboring agents and the corresponding entries in this eigenvector is also established, enabling a distributed neighbor selection algorithm for second-order multi-agent networks. The main results in this paper extend our previous work of distributed neighbor selection algorithm design to multi-agent networks with more complicated agent- level dynamics.
Establishing how a set of learners can provide privacy-preserving federated learning in a fully decentralized (peer-to-peer, no coordinator) manner is an open problem. We propose the first privacy-preserving consensus-based algorithm for the distributed learners to achieve decentralized global model aggregation in an environment of high mobility, where participating learners and the communication graph between them may vary during the learning process. In particular, whenever the communication graph changes, the Metropolis-Hastings method [69] is applied to update the weighted adjacency matrix based on the current communication topology. In addition, the Shamir's secret sharing (SSS) scheme [61] is integrated to facilitate privacy in reaching consensus of the global model. The article establishes the correctness and privacy properties of the proposed algorithm. The computational efficiency is evaluated by a simulation built on a federated learning framework with a real-world dataset.
This article investigates the synthesis of distributed economic control algorithms under which dynamically coupled physical systems are regulated to a variational equilibrium of a constrained convex game. We study two complementary cases: 1) each subsystem is linear and controllable and 2) each subsystem is nonlinear and in the strict-feedback form. The convergence of the proposed algorithms is guaranteed using the Lyapunov analysis. Their performance is verified by two case studies on a multizone building temperature regulation problem and an optimal power flow problem, respectively.
Achieving average consensus without disclosing sensitive information can be a critical concern for multi-agent coordination. This paper examines privacy-preserving average consensus (PPAC) for vector-valued multi-agent networks. In particular, a set of agents with vector-valued states aim to collaboratively reach an exact average consensus of their initial states, while each agent's initial state cannot be disclosed to other agents. We show that the vector-valued PPAC problem can be solved via associated matrix-weighted networks with the higher-dimensional agent state. Specifically, a novel distributed vector-valued PPAC algorithm is proposed by lifting the agent-state to higher-dimensional space and designing the associated matrix-weighted network with dynamic, low-rank, positive semi-definite coupling matrices to both conceal the vector-valued agent state and guarantee that the multi-agent network asymptotically converges to the average consensus. Essentially, the convergence analysis can be transformed into the average consensus problem on switching matrix-weighted networks. We show that the exact average consensus can be guaranteed and the initial agents' states can be kept private if each agent has at least one "legitimate" neighbor. The algorithm, involving only basic matrix operations, is computationally more efficient than cryptography-based approaches and can be implemented in a fully distributed manner without relying on a third party. Numerical simulation is provided to illustrate the effectiveness of the proposed algorithm.
This paper studies optimal motion planning of multiple mobile robots with collision avoidance. We develop a distributed reinforcement learning algorithm which ensures suboptimal goal reaching and anytime collision avoidance simultaneously. Theoretical results on the convergence of neural network weights, the uniform and ultimate boundedness of system states of the closed-loop system, and anytime collision avoidance are established. Numerical simulations for single integrator and unicycle robots illustrate the effectiveness of our theoretical results.
Distributed data sharing in dynamic networks is ubiquitous. It raises the concern that the private information of dynamic networks could be leaked when data receivers are malicious or communication channels are insecure. In this paper, we propose to intentionally perturb the inputs and outputs of a linear dynamic system to protect the privacy of target initial states and inputs from released outputs. We formulate the problem of perturbation design as an optimization problem which minimizes the cost caused by the added perturbations while maintaining system controllability and ensuring the privacy. We analyze the computational complexity of the formulated optimization problem. To minimize the ℓ0 and ℓ2 norms of the added perturbations, we derive their convex relaxations which can be efficiently solved. The efficacy of the proposed techniques is verified by a case study on a heating, ventilation, and air conditioning system.
In this paper, the privacy issue of the recently proposed transactive energy system for electric power system is investigated for the first time. It is identified that the private information of individual market participants will be subject to the risk of leakage during the market interactions. In order to enable the feature of privacy preservance for market participants, a homomorphic encryption-based approach is developed to augment the existing design of transactive energy system. The proposed privacy-preserving design based on the Paillier encryption scheme is then demonstrated on a transactive energy system that coordinates and controls residential air conditioners under the same feeder to manage the feeder congestion. The simulation results confirm the effectiveness of the proposed design in protecting the privacy of individual market participants without affecting the overall system performance.
This paper considers a class of nonlinear distributed control systems subject to false data injection attacks, Byzantine attacks and switching attacks. The problem of attack detection is formulated as the simultaneous recovery of system states, attack vectors and system mode of a switched nonlinear system. In the proposed attack detection algorithm, the inverse system of each mode aims to estimate system states and attack vectors when the corresponding mode is input-output decoupled. A set-valued mode index map gives all modes which generate a switch-singular pair. A machine learning example is used to validate the performance of the developed algorithm.
In this paper, the privacy and security issues associated with transactive energy systems over insecure communications are addressed. In particular, it is ensured that, during market-based interactions: (1) each agent's bidding information remains private; and (2) any extraneous data injection attack can be easily detected. A unified cryptography-based approach that can simultaneously achieve both objectives is developed, where privacy preservation is realized by the Paillier encryption scheme, and attack detection is achieved by the Paillier digital signature scheme. Simulation results verify the effectiveness of the proposed cyber-resilient design for transactive energy systems.
A cyber-physical system (CPS) consists of a large number of geographically dispersed entities and distributed data sharing is necessary to achieve network-wide goals. However, distributed data sharing also raises the significant concern that private or confidential information of legitimate entities could be leaked to unauthorized entities. Privacy has become an issue of high priority to address before a certain CPS can be widely deployed. Existing privacy-preserving techniques solely focus on the cyber space but ignore the physical world, and hence they alone may not be adequate to ensure CPS privacy. This paper aims to summarize recent studies on how to develop control-theoretic approaches to complement existing privacy-preserving techniques and ensure CPS privacy.
This paper studies how a system operator and a set of agents securely execute a distributed projected gradient-based algorithm. In particular, each participant holds a set of problem coefficients and/or states whose values are private to the data owner. The concerned problem raises two questions: how to securely compute given functions; and which functions should be computed in the first place. For the first question, by using the techniques of homomorphic encryption, we propose novel algorithms which can achieve secure multiparty computation with perfect correctness. For the second question, we identify a class of functions which can be securely computed. The correctness and computational efficiency of the proposed algorithms are verified by two case studies of power systems, one on a demand response problem and the other on an optimal power flow problem.
This paper investigates distributed optimization of dynamically coupled networks. We propose distributed algorithms to address two complementary cases: (i) each subsystem is linear and controllable; and (ii) each subsystem is nonlinear and in the strict-feedback form. The convergence of the proposed algorithms is guaranteed using Lyapunov analysis. Their performance is verified by two case studies on an optimal power flow problem and a multizone building temperature regulation problem, respectively.
Suffering from the big "hit" by the Heartbleed attack, the society has learned one hard lesson, namely, the severity of zero-day continuous buffer over-read attacks. According to a survey on Heartbleed, 24-55% of HTTPS servers in the Alexa Top 1 Million were initially vulnerable to Heartbleed, including 44 of the Alexa Top 100. The Heartbleed attack is continuous buffer over-read: it usually lasts several hours, involving hundreds of thousands of probing (buffer over-read) requests. In most cases, a short period of time is insufficient for the attacker to achieve his/her goal. This paper presents our recent work on the development of adaptive defense systems which can practically defend against zero-day continuous buffer over-read attacks; i.e., Heartbleed-like attacks and data structure manipulation attacks, and meanwhile whose cost-effectiveness is mathematically provable.