We present a framework for the automatic configuration of constraints in a disturbed model predictive control (MPC) formulation for multi-agent trajectory generation, leveraging semantic knowledge encoded in a robot world model. Agent routes are planned as sequences of semantic areas to be traversed, enabling the systematic selection of relevant environmental and inter-agent constraints through queries of the world model. The introduction of sub-areas, such as lanes, further reduces computational complexity by limiting the activation of coupling constraints. In addition, the framework supports reasoning over discrete actions, such as yielding to other agents, allowing agents to satisfy high-level task and environmental constraints within the continuous MPC formulation. The proposed framework is validated in both simulation and hardware-in-the-loop experiments using a distributed computing setup. The results demonstrate improved task execution performance while significantly reducing computation time when compared with baseline centralized and distributed formulations.
We present a framework for the automatic configuration of constraints in a disturbed model predictive control (MPC) formulation for multi-agent trajectory generation, leveraging structural knowledge encoded in a robot world model. Agent routes are planned as sequences of areas to be traversed, enabling the systematic selection of relevant environmental and inter-agent constraints through queries of the world model. The introduction of subareas, such as lanes, further reduces computational complexity by limiting the activation of coupling constraints. In addition, the framework supports reasoning over discrete actions, such as yielding to other agents, allowing agents to satisfy high-level task and environmental constraints within the continuous MPC formulation. The proposed framework is validated in both simulation and hardware-in-the-loop experiments using a distributed computing setup. The results demonstrate improved task execution performance while significantly reducing computation time when compared with baseline distributed formulations.
Applying nonlinear model predictive control (NMPC) to systems with hybrid dynamics or discrete actions typically yields mixed-integer nonlinear programs (MINLPs), whose real-time solution remains a major challenge and limits the applicability of mixed-integer NMPC (MINMPC). This paper proposes a myopic MINMPC framework that incorporates value-function approximation to substantially reduce the online computational burden. Using Bellman's principle of optimality, we shorten the prediction horizon and append a value function learned offline from expert state-action demonstrations via inverse optimization with optimality residual minimization. A central feature is the dual treatment of discrete decisions, whereby integer constraints are relaxed during offline learning to enable KKT-residual-based value function synthesis, while the online controller enforces the true integer constraints to ensure feasibility. The learned value function induces a policy that is approximately policy-consistent with the expert demonstrations. The resulting controller achieves high closed-loop performance with a significantly shorter horizon, enabling real-time MINMPC. The effectiveness of the approach is demonstrated on the Lotka-Volterra fishing problem and a satellite attitude control system with discrete actuators.
We present the Output Sensitivity Modification (OSM) strategy for dynamic optimization and a corresponding Nonlinear Model Predictive Controller (NMPC): OSM-NMPC. The OSM strategy redefines the solutions of a dynamic optimization problem, and characterizes them by a set of halting conditions. These conditions are derived from a conceptual Modified-Gradient Sequential Quadratic Programming (MSQP) algorithm, which applies a modification to the model input-output sensitivity matrix at each step of an SQP algorithm applied to the dynamic optimization problem. Previous work has used sensitivity modifications in linear MPC to incorporate structural preferences such as input-output pairing/decoupling. These methods are limited to the linear case, and only assess performance and constraint violation with respect to the modified model. The OSM-NMPC scheme addresses this by extending the method to the nonlinear case, and evaluates feasibility using the true model. To illustrate the efficacy of OSM-NMPC, we apply the proposed method to a benchmark, unstable, exothermic, jacketed Continuous Stirred Tank Reactor (CSTR) process in a numerical case study, where we decouple the inlet flow rate from the tank temperature. The study also compares the OSM-NMPC controller to three other controllers that attempt to achieve decoupling by modifying the cost function.
This paper introduces a model-free real-time optimization (RTO) framework based on unconstrained Bayesian optimization with embedded constraint control. The main contribution lies in demonstrating how this approach simplifies the black-box optimization problem while ensuring "always-feasible" setpoints, addressing a critical challenge in real-time optimization with unknown cost and constraints. Noting that controlling the constraint does not require detailed process models, the key idea of this paper is to control the constraints to "some" setpoint using simple feedback controllers. Bayesian optimization then computes the optimum setpoint for the constraint controllers. By searching over the setpoints for the constraint controllers, as opposed to searching directly over the RTO degrees of freedom, this paper achieves an inherently safe and practical model-free RTO scheme. In particular, this paper shows that the proposed approach can achieve zero cumulative constraint violation without relying on assumptions about the Gaussian process model used in Bayesian optimization. The effectiveness of the proposed approach is demonstrated on a benchmark Williams-Otto reactor example.
Designing algorithms for space bounded models with restoration requirements on (most of) the space used by the algorithm is an important challenge posed about the catalytic computation model introduced by Buhrman et al. (2014). Motivated by the scenarios where we do not need to restore unless w is useful, we relax the restoration requirement: only when the content of the catalytic tape is w ∈ A ⊆ ^* , the catalytic Turing machine needs to restore w at the end of the computation. We define, (A) to be the class of languages that can be accepted by almost-catalytic Turing machines with respect to A (which we call the catalytic set), that uses at most clog n work space and n^c catalytic space. We prove the following for the almost-catalytic model.
This paper proposes a general-purpose multi-agent Bayesian optimization (MABO) where agents are connected via shared variables or constraints, and each agent’s local cost is unknown. The proposed approach is general-purpose in the sense that it can be used with a broad class of decomposition methods, whereby we augment traditional BO acquisition functions with suitably derived coordinating terms to facilitate coordination among subsystems without sharing local data. Regret analysis is also carried out for the general-purpose MABO framework, which reveals that the cumulative regret of the proposed general-purpose MABO is the sum of individual regrets and is independent of the coordinating terms. This adaptability to different decomposition methods ensures versatility across diverse distributed optimization scenarios. Numerical experiments validate the effectiveness of the proposed MABO framework for different classes of decomposition methods.
Reliable core density control with pellet fueling will be necessary to achieve required fusion power output in future fusion tokamaks such as ITER. The discrete nature of fuel pellets, however, complicates the density profile control problem significantly. As a solution, we propose a predictive density profile controller that considers fuel pellets as discrete actuators, while ensuring operation within prescribed density limits. The model predictive control (MPC) scheme we deploy combines the offset-free method to correct prediction model inaccuracies and our novel modified penalty term homotopy algorithm for real-time MPC (PTH-MPC). To demonstrate density profile control with discrete pellets, we couple the PTH-MPC density controller with JINTRAC integrated simulations of the ITER 15 MA/5.3 T scenario, using HPI2 to model discrete pellet ablation and deposition. We compare the density controller performance in integrated simulations using the Bohm/gyro-Bohm turbulent transport model against integrated simulations using the TGLF turbulent transport model. We highlight the necessity of treating pellets as discrete events for controller performance and for remaining within density limits. We conclude that PTH-MPC is a promising candidate for density profile control with pellets fueling in ITER and other future tokamaks and recommend further improvements using learning-based and robust MPC. We also note the limitations of quasi-linear turbulent transport models in simulations involving discrete pellets.
This paper presents a scenario-based model predictive control (MPC) scheme designed to control an evolving pandemic via non-pharmaceutical intervention (NPIs). The proposed approach combines predictions of possible pandemic evolution to decide on a level of severity of NPIs to be implemented over multiple weeks to maintain hospital pressure below a prescribed threshold, while minimizing their impact on society. Specifically, we first introduce a compartmental model which divides the population into Susceptible, Infected, Detected, Threatened, Healed, and Expired (SIDTHE) subpopulations and describe its positive invariant set. This model is expressive enough to explicitly capture the fraction of hospitalized individuals while preserving parameter identifiability w.r.t. publicly available datasets. Second, we devise a scenario-based MPC scheme with recourse actions that captures potential uncertainty of the model parameters. e.g., due to population behavior or seasonality. Our results show that the scenario-based nature of the proposed controller manages to adequately respond to all scenarios, keeping the hospital pressure at bay also in very challenging situations when conventional MPC methods fail.
Control of the density profile based on pellet fueling for the ITER nuclear fusion tokamak involves a multi-rate nonlinear system with safety-critical constraints, input delays, and discrete actuators with parametric uncertainty. To address this challenging problem, we propose a multi-stage MPC (msMPC) approach to handle uncertainty in the presence of mixed-integer inputs. While the scenario tree of msMPC accounts for uncertainty, it also adds complexity to an already computationally intensive mixed-integer MPC (MI-MPC) problem. To achieve real-time density profile controller with discrete pellets and uncertainty handling, we systematically reduce the problem complexity by (1) reducing the identified prediction model size through dynamic mode decomposition with control, (2) applying principal component analysis to reduce the number of scenarios needed to capture the parametric uncertainty in msMPC, and (3) utilizing the penalty term homotopy for MPC (PTH-MPC) algorithm to reduce the computational burden caused by the presence of mixed-integer inputs. We compare the performance and safety of the msMPC strategy against a nominal MI-MPC in plant simulations, demonstrating the first predictive density control strategy with uncertainty handling, viable for real-time pellet fueling in ITER.
The application of supervised learning techniques in combination with model predictive control (MPC) has recently generated significant interest, particularly in the area of approximate explicit MPC, where function approximators like deep neural networks are used to learn the MPC policy via optimal state-action pairs generated offline. While the aim of approximate explicit MPC is to closely replicate the MPC policy, substituting online optimization with a trained neural network, the performance guarantees that come with solving the online optimization problem are typically lost. This paper considers an alternative strategy, where supervised learning is used to learn the optimal value function offline instead of learning the optimal policy. This can then be used as the cost-to-go function in a myopic MPC with a very short prediction horizon, such that the online computation burden reduces significantly without affecting the controller performance. This approach differs from existing work on value function approximations in the sense that it learns the cost-to-go function by using offline-collected state-value pairs, rather than closed-loop performance data. The cost of generating the state-value pairs used for training is addressed using a sensitivity-based data augmentation scheme.
This article considers the problem of steady-state real-time optimization (RTO) of interconnected systems with a common constraint that couples several units, for example, a shared resource. Such problems are often studied under the context of distributed optimization, where decisions are made locally in each subsystem and are coordinated to optimize the overall performance. Here, we use a distributed feedback-optimizing control framework, where the local systems and the coordinator problems are converted into feedback control problems. This is a powerful scheme that allows us to design feedback control loops, estimate parameters locally, and provide a local fast response, allowing different closed-loop time constants for each local subsystem. This article provides a comparative study of different distributed feedback-optimizing control architectures using two case studies. The first case study considers the problem of demand response (DR) in a residential energy hub powered by a common renewable energy source and compares the different feedback-optimizing control approaches using simulations. The second case study experimentally validates and compares the different approaches using a laboratory-scale experimental rig that emulates a subsea oil production network, where the common resource is the gas lift that must be optimally allocated among the wells.
This paper presents a Bayesian optimization framework for the automatic tuning of shared controllers which are defined as a Model Predictive Control (MPC) problem. The proposed framework includes the design of performance metrics as well as the representation of user inputs for simulation-based optimization. The framework is applied to the optimization of a shared controller for an Image Guided Therapy robot. VR-based user experiments confirm the increase in performance of the automatically tuned MPC shared controller with respect to a hand-tuned baseline version as well as its generalization ability.
Application of nonlinear model predictive control (NMPC) to problems with hybrid dynamical systems, disjoint constraints, or discrete controls often results in mixed-integer formulations with both continuous and discrete decision variables. However, solving mixed-integer nonlinear programming problems (MINLP) in real-time is challenging, which can be a limiting factor in many applications. To address the computational complexity of solving mixed integer nonlinear model predictive control problem in real-time, this paper proposes an approximate mixed integer NMPC formulation based on value function approximation. Leveraging Bellman's principle of optimality, the key idea here is to divide the prediction horizon into two parts, where the optimal value function of the latter part of the prediction horizon is approximated offline using expert demonstrations. Doing so allows us to solve the MINMPC problem with a considerably shorter prediction horizon online, thereby reducing the online computation cost. The paper uses an inverted pendulum example with discrete controls to illustrate this approach.
This paper considers the problem of data generation for MPC policy approximation. Learning an approximate MPC policy from expert demonstrations requires a large data set consisting of optimal state-action pairs, sampled across the feasible state space. Yet, the key challenge of efficiently generating the training samples has not been studied widely. Recently, a sensitivity-based data augmentation framework for MPC policy approximation was proposed, where the parametric sensitivities are exploited to cheaply generate several additional samples from a single offline MPC computation. The error due to augmenting the training data set with inexact samples was shown to increase with the size of the neighborhood around each sample used for data augmentation. Building upon this work, this letter paper presents an improved data augmentation scheme based on predictor-corrector steps that enforces a user-defined level of accuracy, and shows that the error bound of the augmented samples are independent of the size of the neighborhood used for data augmentation.
Despite several developments in artificial pancreas technology, postprandial glycemic regulation remains to be a major challenge for type 1 diabetes management. Typically, the large spike in blood glucose concentration induced by meals require an appropriate dose of bolus insulin. Although matching bolus insulin to carbohydrate intake has been shown to improve glycemic regulation, current state-of-the-art meal bolus calculators depend on patient-specific parameters and/or historical clinical data, which may not be easily available. In this paper, we propose a model-free safe and personalized bolus calculator algorithm that is based on safe contextual Bayesian optimization. The proposed algorithm neither requires any patient-specific parameters, nor historical clinical data. Furthermore, the proposed algorithm focuses on patient safety, and ensures satisfaction of the safety-critical hypoglycemia constraint with high probability. In silico experiments conducted on the 10-adult cohort of the FDA-accepted UVA/Padova T1DM simulator, as well as on a cohort of 50 virtual patients based on the Hovorka T1D model in open-loop mode, demonstrate that our algorithm is able to quickly learn the optimum bolus insulin dose for the announced meals using only the patient's CGM data while ensuring patient safety.
This article presents a distributed model predictive controller with time-varying partitioning based on the augmented Lagrangian alternating direction inexact Newton method (ALADIN). In particular, we address the problem of controlling the temperature of a heat transfer fluid (HTF) in a set of loops of solar parabolic collectors by adjusting its flow rate. The control problem involves a nonlinear prediction model, decoupled inequality constraints, and coupled affine constraints on the system inputs. The application of ALADIN to address such a problem is combined with a dynamic clustering-based partitioning approach that aims at reducing, with minimum performance losses, the number of variables to be coordinated. Numerical results on a 10-loop plant are presented.
There are different approaches to tune a control policy that would result in a desired closed-loop performance. The typical design flow involves tuning the control policies offline using (high-fidelity) simulators until satisfactory performance is achieved. This paper on the other hand, considers the problem of tuning parameterized control policies directly by interacting with the real system that also has safety-critical constraints. We use safe Bayesian optimization using interior-point methods to tune the parameterized control policy online that guarantees constraint satisfaction with high probability. The proposed framework is applied to a personalized artificial pancreas system for type 1 diabetes. The paper shows that the parameterized control policy used for blood glucose regulation can be safely tuned to personalize the controller for each individual patient using our approach and thus improve its performance.
Pellet injection is regarded as the only realistic actuator for core density control in future reactors such as ITER and DEMO. However, a control strategy that can reliably regulate the plasma close to operational limits using multiple pellet injectors is not yet available. In this paper, we present the first integrated model control simulations where a dedicated model-predictive controller is included in JINTRAC. We show that, when continuous actuators are considered, a simple transport model with a steady-state disturbance rejection paradigm is capable of capturing the particle transport dynamics for multiple transport models and scenarios. This in turn allows the model-predictive controller to deal with the uncertainty and minimize the control error given the limited actuation space. Furthermore, we show that for ITER and DEMO relevant pellet sizes, the discrete, nonlinear dynamics of pellet injection will limit the control performance and jeopardize the constraints if not accounted for by the controller. Hence, we conclude that for high-performance control on future reactors, controllers will have to be developed that explicitly deal with the discrete pellet dynamics.