We investigate a novel finite-horizon linear-quadratic (LQ) feedback dynamic potential game with a priori unknown cost matrices played between two players. The cost matrices are revealed to the players sequentially, with the potential for future values to be previewed over a short time window. We propose an algorithm that enables the players to predict and track a feedback Nash equilibrium trajectory, and we measure the quality of their resulting decisions by introducing the concept of price of uncertainty. We show that under the proposed algorithm, the price of uncertainty is bounded by horizon-invariant constants. The constants are the sum of three terms; the first and second terms decay exponentially as the preview window grows, and the third depends on the magnitude of the differences between the cost matrices for each player. Through simulations, we illustrate that the resulting price of uncertainty initially decays at an exponential rate as the preview window lengthens, then remains constant for large time horizons.
This article examines state estimation in discrete-time nonlinear stochastic systems with finite-dimensional states and infinite-dimensional measurements, motivated by real-world applications such as vision-based localization and tracking. We develop an extended Kalman filter (EKF) for real-time state estimation, with the measurement noise modeled as an infinite-dimensional random field. When applied to vision-based state estimation, the measurement Jacobians required to implement the EKF are shown to correspond to image gradients. This result provides a novel system-theoretic justification for the use of image gradients as features for vision-based state estimation, contrasting with their (often heuristic) introduction in many computer-vision pipelines. We demonstrate the practical utility of the EKF on a public real-world dataset involving the localization of an aerial drone using video from a downward-facing monocular camera. The EKF is shown to outperform VINS-MONO, an established visual-inertial odometry algorithm, in some cases achieving mean squared error reductions of up to an order of magnitude.
We introduce a class of partially observed Markov decision processes (POMDPs) with costs that can depend on both the value and (future) uncertainty associated with the initial state. These Initial-State Cost POMDPs (ISC-POMDPs) enable the specification of objectives relative to a priori unknown initial states, which is useful in applications such as robot navigation, controlled sensing, and active perception, that can involve controlling systems to revisit, remain near, or actively infer their initial states. By developing a recursive Bayesian fixed-point smoother to estimate the initial state that resembles the standard recursive Bayesian filter, we show that ISC-POMDPs can be treated as POMDPs with (potentially) belief-dependent costs. We demonstrate the utility of ISC-POMDPs, including their ability to select controls that resolve (future) uncertainty about (past) initial states, in simulation.
Systems equipped with modern sensing modalities such as vision and lidar gain access to increasingly high-dimensional measurements with which to enact estimation and control schemes. In this article, we examine the continuum limit of high-dimensional measurements and analyze state estimation in linear time-invariant systems with infinite-dimensional measurements but finite-dimensional states, both corrupted by additive noise. We propose a linear filter and derive the corresponding optimal gain functional in the sense of the minimum mean square error, analogous to the classic Kalman filter. By modeling the measurement noise as a wide-sense stationary random field, we are able to derive the optimal linear filter explicitly, in contrast to previous derivations of Kalman filters in distributed-parameter settings. Interestingly, we find that we need only impose conditions that are finite-dimensional in nature to ensure that the filter is asymptotically stable. The proposed filter is verified via simulation of a linearized system with a pinhole camera sensor.
Event-based cameras are popular for tracking fast-moving objects due to their high temporal resolution, low latency, and high dynamic range. In this paper, we propose a novel algorithm for tracking event blobs using raw events asynchronously in real time. We introduce the concept of an event blob as a spatio-temporal likelihood of event occurrence where the conditional spatial likelihood is blob-like. Many real-world objects such as car headlights or any quickly moving foreground objects generate event blob data. The proposed algorithm uses a nearest neighbour classifier with a dynamic threshold criteria for data association coupled with an extended Kalman filter to track the event blob state. Our algorithm achieves highly accurate blob tracking, velocity estimation, and shape estimation even under challenging lighting conditions and high-speed motions (> 11000 pixels/s). The microsecond time resolution achieved means that the filter output can be used to derive secondary information such as time-to-contact or range estimation, that will enable applications to real-world problems such as collision avoidance in autonomous driving.
We propose Zeroth-Order Random Matrix Search for Learning from Demonstrations (ZORMS-LfD). ZORMS-LfD enables the costs, constraints, and dynamics of constrained optimal control problems, in both continuous and discrete time, to be learned from expert demonstrations without requiring smoothness of the learning-loss landscape. In contrast, existing state-of-the-art first-order methods require the existence and computation of gradients of the costs, constraints, dynamics, and learning loss with respect to states, controls and/or parameters. Most existing methods are also tailored to discrete time, with constrained problems in continuous time receiving only cursory attention. We demonstrate that ZORMS-LfD matches or surpasses the performance of state-of-the-art methods in terms of both learning loss and compute time across a variety of benchmark problems. On unconstrained continuous-time benchmark problems, ZORMS-LfD achieves similar loss performance to state-of-the-art first-order methods with an over 80% reduction in compute time. On constrained continuous-time benchmark problems where there is no specialized state-of-the-art method, ZORMS-LfD is shown to outperform the commonly used gradient-free Nelder-Mead optimization method. We illustrate the practicality of ZORMS-LfD on a human motion dataset, and derive complexity bounds for it on problems with Lipschitz continuous (but potentially nondifferentiable) loss.
A new surveillance-evasion differential game is posed and solved in which an agile pursuer (the prying pedestrian) seeks to remain within a given surveillance range of a less agile evader that aims to escape. In contrast to previous surveillance-evasion games, the pursuer is agile in the sense of being able to instantaneously change the direction of its velocity vector, while the evader is constrained to have a finite maximum turn rate. Both the game of kind, concerned with conditions under which the evader can escape, and the game of degree, concerned with the evader seeking to minimize the escape time while the pursuer seeks to maximize it, are considered. The game-of-degree solution is surprisingly complex compared to solutions to analogous pursuit-evasion games with an agile pursuer because it exhibits dependence on the ratio of the pursuer’s speed to the evader’s speed. It is, however, surprisingly simple compared to solutions to classic surveillance-evasion games with a turn-limited pursuer.
The trade-off between communication resources and estimation accuracy is widely considered in sensor networks. In this paper we consider the problem of estimating the trajectory of an event-triggered hidden Markov model, where the controller decides at each time step whether or not the sensor should sample and transmit a measurement to the estimator. Adopting a Shannon information-theoretic point of view, we quantify the required communication resources by the entropy of the transmitted observation sequence, with a special symbol to denote non-transmission. Furthermore we evaluate the trajectory uncertainty by the conditional entropy of the state sequence given the received observations. Simultaneous minimization of the communication resources and state uncertainty is formulated and solved within a partially observable Markov decision process framework, yielding a threshold policy for triggering transmissions.
We formulate the discrete-time inverse optimal control problem of inferring unknown parameters in the objective function of an optimal control problem from measurements of optimal states and controls as a nonlinear filtering problem. This formulation enables us to propose a novel extended Kalman filter (EKF) for solving inverse optimal control problems in a computationally efficient recursive online manner that requires only a single pass through the measurement data. Importantly, we show that the Jacobians required to implement our EKF can be computed efficiently by exploiting recent Pontryagin differentiable programming results, and that our consideration of an EKF enables the development of first-of-their-kind theoretical error guarantees for online inverse optimal control with noisy incomplete measurements. Our proposed EKF is shown to be significantly faster than an alternative unscented Kalman filter-based approach.
In this paper, we describe an undesirable weak practical super-martingale hallucination phenomenon that can emerge in the Bayesian quickest detection and identification problem. We establish that when measurements are insufficiently informative, a situation described by a relative entropy condition on measurement densities, the Bayesian quickest detection and identification solution can (undesirably) become increasingly confident that a change has occurred, even when it has not. Finally, we illustrate the phenomenon in simulation studies and the vision-based aircraft detection application which illustrates the optimal rule can be unsuitable in the sense of hallucinating a change that has not occurred.
In this paper we present a novel framework for quickly detecting a change in a general dependent stochastic process. We propose that any general dependent Bayesian quickest change detection (QCD) problem can be converted into a hidden Markov model (HMM) QCD problem, provided that a suitable state process can be constructed. The optimal rule for HMM QCD is then a simple threshold test on the posterior probability of a change. We investigate case studies that can be considered structured generalisations of Bayesian HMM QCD problems including: quickly detecting changes in statistically periodic processes and quickest detection of a moving target in a sensor network. Using our framework we pose and establish the optimal rules for these case studies. We also illustrate the performance of our optimal rule on real air traffic data to verify its simplicity and effectiveness in detecting changes.
We establish the Lipschitz continuity of the value functions of an active fixed-sample-size hypothesis testing problem when it is reformulated as a partially observed Markov decision process. These Lipschitz results enable us to develop novel upper and lower bounds on the value of information, which is the expected difference between the value functions before and after performing an experiment. Our novel Lipschitz and value-of-information results provide new practical insight into optimal policies for active fixed-sample-size hypothesis testing without resorting to approximate dynamic programming schemes or asymptotic analysis with infinite numbers of samples. We illustrate the utility of our results by showing that a simple scheme based on selecting experiments that maximize a value-of-information bound achieves near-optimal performance in simulations.
We investigate partially observed Markov decision processes (POMDPs) with cost functions regularized by entropy terms describing state, observation, and control uncertainty. Standard POMDP techniques are shown to offer bounded-error solutions to these entropy-regularized POMDPs, with exact solutions possible when the regularization involves the joint entropy of the state, observation, and control trajectories. Our joint-entropy result is particularly surprising since it constitutes a novel, tractable formulation of active state estimation.
In this paper, we propose and analyze a new method for online linear quadratic regulator (LQR) control with a priori unknown time-varying cost matrices. The cost matrices are revealed sequentially with the potential for future values to be previewed over a short window. Our novel method involves using the available cost matrices to predict the optimal trajectory, and a tracking controller to drive the system towards it. We adopted the notion of dynamic regret to measure the performance of this proposed online LQR control method, with our main result being that the (dynamic) regret of our method is upper bounded by a constant. Moreover, the regret upper bound decays exponentially with the preview window length, and is extendable to systems with disturbances. We show in simulations that our proposed method offers improved performance compared to other previously proposed online LQR methods.
Recent empirical success has led to a rise in popularity of the options framework for Hierarchical Reinforcement Learning (HRL). This framework tackles the scalability problem in Reinforcement Learning (RL) by introducing a layer of abstraction (i.e. high-level options) over the (low-level) decision process. Hierarchical Imitation Learning (HIL) is the problem of learning low-level and high-level policies within HRL from expert demonstrations consisting only of the low-level actions and states, with the high-level options being hidden (or latent). Due to the latent options, recent work on HIL has focused on the development of Expectation-Maximization (EM) algorithms inspired by approaches such as the celebrated Baum-Welch algorithm for hidden Markov models (HMMs). In this work, we take a different approach and derive a new HIL framework inspired by the spectral method of moments for HMMs. The method of moments offers global and consistent convergence under mild regulatory conditions, whilst only requiring one sweep through the data set of state and action pairs, giving it a competitive run time.
We investigate the problem of finding paths that enable a robot modeled as a Dubins car (i.e., a constant-speed finite-turn-rate unicycle) to escape from a circular region of space in minimum time. This minimum-time escape problem arises in marine, aerial, and ground robotics in situations where a safety region has been violated and must be exited before a potential negative consequence occurs (e.g., a collision). Using the tools of nonlinear optimal control theory, we show that a surprisingly simple closed-form feedback control law solves this minimum-time escape problem, and that the minimum-time paths have an elegant geometric interpretation.
This letter establishes that an exactly optimal rule for Bayesian Quickest Change Detection (QCD) of Markov chains is an optimal stopping rule in the form of a threshold test on the no change posterior. We also provide a computationally efficient scalar filter for the no change posterior whose effort is independent of the dimension of the chains. We establish that an (undesirable) weak practical super-martingale phenomenon can be exhibited by the no change posterior when the before and after chains are too close in a relative entropy rate sense. The proposed detector is examined in simulation studies.
This work proposes a novel approach to approximate optimal linear filters for discrete-time linear Gaussian systems with infinite-dimensional measurements and finite- dimensional states. Assuming scalar-valued states for simplicity, we formulate the problem in terms of optimally selecting $N$ points at which to sample the infinite-dimensional measurement, in order to minimize the mean-squared filtering error. We show that for large N, this problem can be expressed using the notion of an asymptotic point density function from the field of high-resolution quantization theory. To the best of the authors' knowledge, this method has not been considered in infinite- dimensional filtering previously. This leads to a characterization in terms of an Urysohn integral equation, which can be solved numerically to yield an asymptotically optimal $N$ -point filter. The mean-squared approximation error is proportional to N -4 , which is faster than the typical N -2 decay of high-resolution quantization and suggests that this approximation method will be useful even for moderate or small N. These properties are verified by simulations based on a linearized pinhole camera measurement model.
This paper proposes a new method for differentiating through optimal trajectories arising from non-convex, constrained discrete-time optimal control (COC) problems using the implicit function theorem (IFT). Previous works solve a differential Karush-Kuhn-Tucker (KKT) system for the trajectory derivative, and achieve this efficiently by solving an auxiliary Linear Quadratic Regulator (LQR) problem. In contrast, we directly evaluate the matrix equations which arise from applying variable elimination on the Lagrange multiplier terms in the (differential) KKT system. By appropriately accounting for the structure of the terms within the resulting equations, we show that the trajectory derivatives scale linearly with the number of timesteps. Furthermore, our approach allows for easy parallelization, significantly improved scalability with model size, direct computation of vector-Jacobian products and improved numerical stability compared to prior works. As an additional contribution, we unify prior works, addressing claims that computing trajectory derivatives using IFT scales quadratically with the number of timesteps. We evaluate our method on a both synthetic benchmark and four challenging, learning from demonstration benchmarks including a 6-DoF maneuvering quadrotor and 6-DoF rocket powered landing.
Event-based cameras have become increasingly popular for tracking fast-moving objects due to their high temporal resolution, low latency, and high dynamic range. In this paper, we propose a novel algorithm for tracking event blobs using raw events asynchronously in real time. We introduce the concept of an event blob as a spatio-temporal likelihood of event occurrence where the conditional spatial likelihood is blob-like. Many real-world objects generate event blob data, for example, flickering LEDs such as car headlights or any small foreground object moving against a static or slowly varying background. The proposed algorithm uses a nearest neighbour classifier with a dynamic threshold criteria for data association coupled with a Kalman filter to track the event blob state. Our algorithm achieves highly accurate tracking and event blob shape estimation even under challenging lighting conditions and high-speed motions. The microsecond time resolution achieved means that the filter output can be used to derive secondary information such as time-to-contact or range estimation, that will enable applications to real-world problems such as collision avoidance in autonomous driving.