In this tutorial, we explore the intersection of quantum control, game theory, and artificial intelligence, emphasizing their combined potential to enhance secure and strategic decision-making in complex systems. It examines how deceptive behavior can be understood and influenced within strategic frameworks, and outlines principles for designing resilient systems that leverage advantages in both quantum and learning-driven environments. While the discussion centers on specific examples, the foundational ideas are broadly applicable across a range of domains.
Accurately modeling friction in robotics remains a core challenge, as robotics simulators like MuJoCo and PyBullet use simplified friction models or heuristics to balance computational efficiency with accuracy, where these simplifications and approximations can lead to substantial differences between simulated and physical performance. In this paper, we present a physics-informed friction estimation framework that enables the integration of well-established friction models with learnable components, requiring only minimal, generic measurement data. Our approach enforces physical consistency yet retains the flexibility to capture complex friction phenomena. We demonstrate, on an underactuated and nonlinear system, that the learned friction models, trained solely on small and noisy datasets, accurately reproduce dynamic friction properties with significantly higher fidelity than the simplified models commonly used in robotics simulators. Crucially, we show that our approach enables the learned models to be transferable to systems they are not trained on. This ability to generalize across multiple systems streamlines friction modeling for complex, underactuated tasks, offering a scalable and interpretable path toward improving friction model accuracy in robotics and control.
The recent availability of a technology providing real-time, seconds-resolved in vivo drug concentration measurements has opened the door to performing fully autonomous, closed-loop feedback control over drug dosing. The controllers employed in prior demonstrations of such dosing, however, were designed and optimized using population-based pharmacokinetic models. In the face of individual pharmacokinetic variation (between subjects or even within a single subject over time as their physiology varies), these controllers must be set rather conservatively so as to avoid potentially dangerous overshooting. This, however, slows the speed with which they achieve the desired set point and opening the possibility of their nevertheless still overshooting if the response of a subject differs too much from that of the "average" subject. To address these issues, here we have developed an adaptive feedback-control system that, rather than employing population-pharmacokinetic information, instead uses real-time drug concentration measurements to individualize drug delivery to the specific subject, "on the fly." To achieve this, the system estimates the pharmacokinetics of the individual subject during the initial stages of the infusion, and then continuously updates this subject-specific pharmacokinetic model to maintain effective controller performance even in the face of physiological variations brought on, for example, by changing health status. Using this approach, we then demonstrate a feedback controller that rapidly (20-30 min) achieves and accurately (5%-12% root-mean-squared deviations, though this also includes sensor noise) maintains pre-defined concentrations of the anesthetic procaine in the ventricles of live rats, an application that, due to the delays associated with intracranial drug transport, represents a particular challenge for feedback-controlled intravenous drug delivery. Given the precision and accuracy it achieves in our rat animal model, we believe that the use of adaptive feedback control will ultimately enable safer, more precise drug dosing in humans.
Multi-robot systems are well-positioned for exploration in hazardous environments, but effective deployment requires deciding not only where robots should gather information, but also how risk should be distributed across heterogeneous team members. This paper develops a game-theoretic framework for cooperative risk-aware exploration based on ecologically inspired altruistic behavior. Each robot selects a finite-horizon trajectory to maximize information gain while penalizing redundant exploration and expected hazard exposure. Heterogeneity is introduced through agent-specific value parameters for encoding altruistic coupling, which is modeled through relatedness weights inspired by Hamilton's rule. We introduce a game-theoretic structure for trajectory planning that defines a Social Nash Equilibrium, which modifies the utility of agent actions according to agent relatedness. This utility shaping causes agents to internalize the effect of their trajectory choices on teammates, encouraging lower-valued robots to accept risk when doing so benefits higher-valued agents and improves team performance. We define an exploration utility for agents that rewards area coverage and uncertainty reduction, while also penalizing redundancy and risk, enabling projected gradient-based waypoint optimization in a receding-horizon planner. Simulations show that altruistic planning reduces redundant exploration, improves inter-robot separation, and reallocates risk according to agent value while maintaining comparable map coverage. We further demonstrate the approach in hardware experiments, where planned waypoints are tracked by wheeled robots using single-integrator controllers and barrier certificates.
We consider a discrete-time event-triggered control setting in which a scheduler collocated with the plant's sensors decides when to transmit the full-state to a remote controller collocated with the actuators. The scheduler can only use state information available since the previous transmission time. If the scheduler transmits periodically with a period larger than or equal to one, the $H_\infty$ optimal controller guarantees an optimal attenuation bound ($\ell _{2}$ gain) from any square summable disturbance input to a plant's output. We show that, for the considered setting and under additional assumptions, there does not exist a controller and scheduler pair that strictly improves the optimal attenuation bound of periodic control while achieving a lower average transmission rate. Equivalently, given any controller and scheduler pair, there exists a square summable disturbance such that either the attenuation bound or the average transmission rate is larger than or equal to that of optimal periodic control.
The Koopman operator, which describes a dynamical system via a linear representation that can be approximately learned from data, allows application of linear control techniques to nonlinear systems. This work considers a nonlinear optimal control problem with final state constraints, which we alternatively represent by a linear problem on a lifted state space described by the Koopman operator. We show that the optimal cost-to-go to this linear problem has a piecewise linear form with respect to the lifted state. A robust formulation, in which we minimize the worst-case cost with respect to bounded errors in the initial state and Koopman dynamics, retains a piecewise affine structure. Due to the combinatorial nature of the problem, the number of possible input sequences grows exponentially in the horizon, and so we provide a heuristic pruning algorithm that reduces the search space to a much smaller subset. Our control approach is generally applicable, but its effectiveness depends on a domain-specific choice of lifting functions, termed observables. For a class of systems with dynamics described by some number of distinct “entities”, we define observables that describe “densities” of these entities, as well as products of densities that capture a richer class of interactions between entities.
This study presents the optimal solutions to the state feedback finite horizon disturbance attenuation problem using a game theory approach. The proposed solution method is valid for linear dynamical systems and quadratic objective functions under bounded disturbances. Two solution regions in the space of initial states define the optimal value of the controller. The first region, which contains the zero initial state, features the linear optimal H-infinity controller, whereas the second region is characterized by a nonlinear optimal control, which converges to the linear quadratic regulator (LQR) for large initial states. The transition between the two regions provides a unified framework that spans from H-infinity control to LQR control, depending on the relative sizes of the initial state and the bound on the disturbance. A novel solution algorithm is introduced, reducing the optimization of the disturbance attenuation problem to a linear algebra problem for the first region and a single, scalar nonlinear program for the second region. The performance of the algorithm is tested on a representative set of numerical examples. This article enhances the versatility of disturbance attenuation state feedback controllers by introducing an efficient optimal solution strategy applicable to both zero and nonzero initial states. One consequence of these results is that optimal disturbance attenuation of even linear systems requires nonlinear control.
Electrochemical aptamer-based (EAB) sensors enable the continuous, real-time monitoring of drugs and biomarkers in situ in the blood, brain, and peripheral tissues of live subjects. The real-time concentration information produced by these sensors provides unique opportunities to perform closed-loop, feedback-controlled drug delivery, by which the plasma concentration of a drug can be held constant or made to follow a specific, time-varying profile. Motivated by the observation that the site of action of many drugs is the solid tissues and not the blood, here we experimentally confirm that maintaining constant plasma drug concentrations also produces constant concentrations in the interstitial fluid (ISF). Using an intravenous EAB sensor we performed feedback control over the concentration of doxorubicin, an anthracycline chemotherapeutic, in the plasma of live rats. Using a second sensor placed in the subcutaneous space, we find drug concentrations in the ISF rapidly (30-60 min) match and then accurately (RMS deviation of 8 to 21%) remain at the feedback-controlled plasma concentration, validating the use of feedback-controlled plasma drug concentrations to control drug concentrations in the solid tissues that are the site of drug action. We expanded to pairs of sensors in the ISF, the outputs of the individual sensors track one another with good precision (R 2 = 0.95-0.99), confirming that the performance of in vivo EAB sensors matches that of prior, in vitro validation studies. These observations suggest EAB sensors could prove a powerful new approach to the high-precision personalization of drug dosing.
This paper deals with the problem of deciding which inputs or parameters in a dynamic model have the greatest impact on the state or output of a linear system. In over-parameterised models or Multi-Agent Systems (MAS), a sensitivity analysis tool would allow for reduced representations of the same system or to decide which agents should be monitored. The proposed approach consists in constructing the exact reachable set for a linear system (uncertain in the case of parameter sensitivity) and projecting onto the desired coordinates associated with the state or output in question. Given the exact representation, under mild assumptions, sensitivity values are optimal for the proposed functions, although others can be specified using the same sets. Illustrative examples are presented in order to provide insights on the proposed method.
We consider the problem of robust control of an unknown but minimal linear time-invariant system under input and output disturbances and adversarial manipulation. An attacker can (i) corrupt sensor measurements (deception attacks) and (ii) perturb the control channel (actuation attacks). To address the lack of model knowledge, we adapt the Data-enabled Predictive Control (DeePC) framework, which constructs predictors directly from input–output data with bounded disturbance. We formulate a finite-horizon open-loop control problem as a two-player zero-sum game with asymmetric information: the defender selects control inputs based on measured data and only knows an upper bound on disturbances, whereas the attacker has access to the true disturbance realization and can remain stealthy by hiding within this uncertainty set. The main contributions are (i) sufficient conditions for the existence of a Nash equilibrium corresponding to saddle-point policies for this game, and (ii) an analysis of the defender’s security strategy against deception and actuation attacks. Simulation studies on finite-horizon control demonstrate the effectiveness of the proposed approach.
This paper addresses zero-sum “turn” games, in which only one player can make decisions at each state. We show that pure saddle-point state-feedback policies for turn games can be constructed from dynamic programming fixed-point equations for a single value function or Q-function. These fixed-points can be constructed using a suitable form of Q-learning. For discounted costs, convergence of this form of Q-learning can be established using classical techniques. For undiscounted costs, we provide a convergence result that applies to finite-time deterministic games, which we use to illustrate our results. For complex games, the Q-learning iteration must be terminated before exploring the full-state, which can lead to policies that cannot guarantee the security levels implied by the final Q-function. To mitigate this, we propose an “opponent-informed” exploration policy for selecting the Q-learning samples. This form of exploration can guarantee that the final Q-function provides security levels that hold, at least, against a given set of policies. A numerical demonstration for a multi-agent game, Atlatl, indicates the effectiveness of these methods.
Background and PurposeThe ability to measure specific molecules at multiple sites within the body simultaneously, and with a time resolution of seconds, could greatly advance our understanding of drug transport and elimination.Experimental ApproachAs a proof‐of‐principle demonstration, here we describe the use of electrochemical aptamer‐based (EAB) sensors to measure transport of the antibiotic vancomycin from the plasma (measured in the jugular vein) to the cerebrospinal fluid (measured in the lateral ventricle) of live rats with temporal resolution of a few seconds.Key ResultsIn our first efforts, we made measurements solely in the ventricle. Doing so we find that, although the collection of hundreds of concentration values over a single drug lifetime enables high‐precision estimates of the parameters describing intracranial transport, due to a mathematical equivalence, the data produce two divergent descriptions of the drug's plasma pharmacokinetics that fit the in‐brain observations equally well. The simultaneous collection of intravenous measurements, however, resolves this ambiguity, enabling high‐precision (typically of ±5 to ±20% at 95% confidence levels) estimates of the key pharmacokinetic parameters describing transport from the blood to the cerebrospinal fluid in individual animals.Conclusions and ImplicationsThe availability of simultaneous, high‐density ‘in‐vein’ (plasma) and ‘in‐brain’ (cerebrospinal fluid) measurements provides unique opportunities to explore the assumptions almost universally employed in earlier compartmental models of drug transport, allowing the quantitative assessment of, for example, the pharmacokinetic effects of physiological processes such as the bulk transport of the drug out of the CNS via the dural venous sinuses.
We propose a Markov Chain Monte Carlo (MCMC) algorithm based on Gibbs sampling with parallel tempering to solve nonlinear optimal control problems. The algorithm is applicable to nonlinear systems with dynamics that can be approximately represented by a finite dimensional Koopman model, potentially with high dimension. This algorithm exploits linearity of the Koopman representation to achieve significant computational saving for large lifted states. We use a video-game to illustrate the use of the method.
Data-driven control benefits from rich datasets, but constructing such datasets becomes challenging when gathering data is limited. We consider an offline experiment design approach to gathering data where we design a control input to collect data that will most improve the performance of a feedback controller. We show how such a control-oriented approach can be used in a setting with linear dynamics and quadratic objective and, through design of a gradient estimator, solve the problem via stochastic gradient descent. We contrast our method with a classical A-optimal experiment design approach and numerically demonstrate that our method outperforms A-optimal design in terms of improving control performance.
Data-driven control benefits from rich datasets, but constructing such datasets becomes challenging when gathering data is limited. We consider an offline experiment design approach to gathering data where we design a control input to collect data that will most improve the performance of a feedback controller. We consider a setting in which the dynamics are modeled parametrically and formulate a control-oriented identification procedure by way of a stochastic optimization problem that explicitly optimizes the post-experiment closed-loop control performance. We propose solving this problem via stochastic gradient descent by first constructing a gradient estimator of our stochastic objective. We then focus on a particular setting with linear dynamics and quadratic objective, which benefits from a numerically tractable gradient estimator. We show our formulation numerically outperforms an A-and L-optimal experiment design approach, illustrate the effects of scaling the system state and input dimensions, and compare against a recent robust dual control approach.
This paper presents an in-reachability based classification of invariant synchrony patterns in coupled cell networks (CCNs). These patterns are encoded through partitions on the set of cells, whose subsets of synchronised cells are called colours. We study the influence of the structure of the network in the qualitative behaviour of invariant synchrony sets, in particular, with respect to the different types of (cumulative) in-neighbourhoods and the in-reachability sets. This motivates the proposed approach to classify the partitions into the categories of strong, rooted and weak, according to how their colours are related with respect to the connectivity structure of the network. Furthermore, we show how this classification system acts under the partition join ( ∨ ) operation, which gives us the synchrony pattern that corresponds to the intersection of synchrony sets.
We consider a discrete-time linear system for which the control input is updated at every sampling time, but the state is measured at a slower rate. We allow the state to be sampled according to a periodic schedule, which dictates when the state should be sampled along a period. Given a desired average sampling interval, our goal is to determine sampling schedules that are optimal in the sense that they minimize the $h_2$ or the $h_\infty$ closed-loop norm, under an optimal state-feedback control law. Our results show that, when the desired average sampling interval is an integer, the optimal state sampling turns out to be evenly spaced. This result indicates that, for the $h_2$ and $h_\infty$ performance metrics, there is relatively little benefit to go beyond constant-period sampling.
In this article, we address the problem of minimizing an expected value with stochastic constraints, known in the literature as stochastic programming. Our approach is based on computing and optimizing bounds for the expected value that are obtained by solving a deterministic optimization problem that uses the probability density function (pdf) to penalize unlikely values for the random variables. The suboptimal solution obtained through this approach has performances guarantees with respect to the optimal one, while satisfying stochastic and deterministic constraints. We illustrate this approach in the context of the following three different classes of optimization problems: finite horizon optimal stochastic control, with state or output feedback; parameter estimation with latent variables; and nonlinear Bayesian experiment design. By the means of several numerical examples, we show that our suboptimal solution achieves results similar to those obtained with Monte Carlo methods with a fraction of the computational burden, highlighting the usefulness of this approach in real-time optimization problems.
Learning for control in repeated tasks allows for well-designed experiments to gather the most useful data. We consider the setting in which we use a data-driven controller that does not have access to the true system dynamics. Rather, the controller uses inferred dynamics based on the available information. In order to acquire data that is beneficial for this controller, we present an experimental design approach that leverages the current data to improve expected control performance. We focus on the setting in which inference on the unknown dynamics is performed using Gaussian processes. Gaussian processes not only provide uncertainty quantification but also allow us to leverage structures inherit to Gaussian random variables. Through this structure, we design experiments via gradient descent on the expected control performance with respect to the experiment input. In particular, we focus on a chance-constrained minimum expected time control problem. Numerical demonstrations of our approach indicate our experimental design outperforms relevant benchmarks.
We introduce a performance-guaranteed limbic system-inspired control (LISIC) strategy for nonlinear multi-agent systems (MASs) with uncertain high-order dynamics and external perturbations, where each agent in the MAS incorporates a LISIC structure to support the consensus controller. This novel approach, which we call double integrator LISIC (DILISIC), is designed to imitate double integrator dynamics after closing the agent-specific control loop, allowing the control designer to apply consensus techniques specifically formulated for double integrator agents. The objective of each DILISIC structure is then to identify and compensate model differences between the theoretical assumptions considered when tuning the consensus protocol and the actual conditions encountered in the real-time system to be controlled. A Lyapunov analysis is provided to demonstrate the stability of the closed-loop MAS enhanced with the DILISIC. Additionally, the stabilization of a complex system via DILISIC is addressed in a synthetic scenario: the consensus control of a team of flexible single-link arms. The dynamics of these agents are of fourth order, contain uncertainties, and are subject to external perturbations. The numerical results validate the applicability of the proposed method.
A Stephen Morse合作论文数Department of Electrical Engineering, School of Engineering and Applied Science, Yale University25
Maria Prandini合作论文数Dipartimento di Elettronica, Informazione e Bioingegneria, Politecnico di Milano9