Consider a linear stochastic filtering problem in which the probability measure specifying all randomness is only partially known. The deviation between the real and assumed probability models is constrained by a divergence bound between the respective probability measures under which the models are defined. This bound defines a so-called uncertainty set. A recursive set-valued filtering characterization is derived and is guaranteed (with probability one) to contain the true conditional posterior of the unknown, real world, filtering problem when the real world measure is within this uncertainty set. Some filtering approximations and related results are given. The set-valued characterization is related to the problem of robust model validation and model goodness-of-fit statistical hypothesis testing. It is shown how relevant terms involving the innovation sequence (re)appear in multiple settings from set-valued filtering to statistical model evaluation.
The purpose of this review is to present a comprehensive overview of the theory of ensemble Kalman-Bucy filtering for continuous-time, linear-Gaussian signal and observation models. We present a system of equations that describe the flow of individual particles and the flow of the sample covariance and the sample mean in continuous-time ensemble filtering. We consider these equations and their characteristics in a number of popular ensemble Kalman filtering variants. Given these equations, we study their asymptotic convergence to the optimal Bayesian filter. We also study in detail some non-asymptotic time-uniform fluctuation, stability, and contraction results on the sample covariance and sample mean (or sample error track). We focus on testable signal/observation model conditions, and we accommodate fully unstable (latent) signal models. We discuss the relevance and importance of these results in characterising the filter's behaviour, e.g. it's signal tracking performance, and we contrast these results with those in classical studies of stability in Kalman-Bucy filtering.We also provide a novel (and negative) result proving that the bootstrap particle filter cannot track even the most basic unstable latent signal, in contrast with the ensemble Kalman filter (and the optimal filter). We provide intuition for how the main results extend to nonlinear signal models and comment on their consequence on some typical filter behaviours seen in practice, e.g. catastrophic divergence.
We consider the Bayesian optimal filtering problem: i.e. estimating some conditional statistics of a latent time-series signal from an observation sequence. Classical approaches often rely on the use of assumed or estimated transition and observation models. Instead, we formulate a generic recurrent neural network framework and seek to learn directly a recursive mapping from observational inputs to the desired estimator statistics. The main focus of this article is the approximation capabilities of this framework. We provide approximation error bounds for filtering in general non-compact domains. We also consider strong time-uniform approximation error bounds that guarantee good long-time performance. We discuss and illustrate a number of practical concerns and implications of these results.
As a fundamental structure in real-world networks, in addition to graph topology, communities can also be reflected by abundant node attributes. In attributed community detection, probabilistic generative models (PGMs) have become the mainstream method due to their principled characterization and competitive performances. Here, we propose a novel PGM without imposing any distributional assumptions on attributes, which is superior to the existing PGMs that require attributes to be categorical or Gaussian distributed. Based on the block model of graph structure, our model incorporates the attribute by describing its effect on node popularity. To characterize the effect quantitatively, we analyze the community detectability for our model and then establish the requirements of the node popularity term. This leads to a new scheme for the crucial model selection problem in choosing and solving attributed community detection models. With the model determined, an efficient algorithm is developed to estimate the parameters and to infer the communities. The proposed method is validated from two aspects. First, the effectiveness of our algorithm is theoretically guaranteed by the detectability condition. Second, extensive experiments indicate that our method not only outperforms the competing approaches on the employed datasets, but also shows better applicability to networks with various node attributes.
The capability of recurrent neural networks to approximate trajectories of a random dynamical system, with random inputs, on non-compact domains, and over an indefinite or infinite time horizon is considered. The main result states that certain random trajectories over an infinite time horizon may be approximated to any desired accuracy, uniformly in time, by a certain class of deep recurrent neural networks, with simple feedback structures. The formulation here contrasts with related literature on this topic, much of which is restricted to compact state spaces and finite time intervals. The model conditions required here are natural, mild, and easy to test, and the proof is very simple.
The article presents a rather surprising Floquet-type representation of time-varying transition matrices associated with a class of nonlinear matrix differential Riccati equations. The main difference with conventional Floquet theory comes from the fact that the underlying flow of the solution matrix is aperiodic. The monodromy matrix associated with this Floquet representation coincides with the exponential (fundamental) matrix associated with the stabilising fixed point of the Riccati equation. The second part of this article is dedicated to the application of this representation to the stability of matrix differential Riccati equations. We provide refined global and local contraction inequalities for the Riccati exponential semigroup that depend linearly on the spectral norm of the initial condition. These refinements improve upon existing results and are a direct consequence of the Floquet-type representation, yielding what seems to be the first results of this type for this class of models.
Distributed consensus in the Wasserstein metric space of probability measures on the real line is introduced in this work. Convergence of each agent's measure to a common measure is proven under a weak network connectivity condition. The common measure reached at each agent is one minimizing a weighted sum of its Wasserstein distance to all initial agent measures. This measure is known as the Wasserstein barycenter. Special cases involving Gaussian measures, empirical measures, and time-invariant network topologies are considered, where convergence rates and average-consensus results are given. This work has possible applicability in computer vision, machine learning, clustering, and estimation.
We present a backward diffusion flow (i.e. a backward-in-time stochastic differential equation) whose marginal distribution at any (earlier) time is equal to the smoothing distribution when the terminal state (at a latter time) is distributed according to the filtering distribution. This is a novel interpretation of the smoothing solution in terms of a nonlinear diffusion (stochastic) flow. This solution contrasts with, and complements, the (backward) deterministic flow of probability distributions (viz. a type of Kushner smoothing equation) studied in a number of prior works. A number of corollaries of our main result are given including a derivation of the time-reversal of a stochastic differential equation, and an immediate derivation of the classical Rauch-Tung-Striebel smoothing equations in the linear setting.
Sequential Monte Carlo methods, also known as particle methods, are a popular set of techniques for approximating high-dimensional probability distributions and their normalizing constants. These methods have found numerous applications in statistics and related fields; e.g. for inference in non-linear non-Gaussian state space models, and in complex static models. Like many Monte Carlo sampling schemes, they rely on proposal distributions which crucially impact their performance. We introduce here a class of controlled sequential Monte Carlo algorithms, where the proposal distributions are determined by approximating the solution to an associated optimal control problem using an iterative scheme. This method builds upon a number of existing algorithms in econometrics, physics, and statistics for inference in state space models, and generalizes these methods so as to accommodate complex static models. We provide a theoretical analysis concerning the fluctuation and stability of this methodology that also provides insight into the properties of related algorithms. We demonstrate significant gains over state-of-the-art methods at a fixed computational complexity on a variety of applications.
Matrix differential Riccati equations are central in filtering and optimal control theory. The purpose of this article is to develop a perturbation theory for a class of stochastic matrix Riccati diffusions. Diffusions of this type arise, for example, in the analysis of ensemble Kalman-Bucy filters since they describe the flow of certain sample covariance estimates. In this context, the random perturbations come from the fluctuations of a mean field particle interpretation of a class of nonlinear diffusions equipped with an interacting sample covariance matrix functional. The main purpose of this article is to derive non-asymptotic Taylor-type expansions of stochastic matrix Riccati flows with respect to some perturbation parameter. These expansions rely on an original combination of stochastic differential analysis and nonlinear semigroup techniques on matrix spaces. The results here quantify the fluctuation of the stochastic flow around the limiting deterministic Riccati equation, at any order. The convergence of the interacting sample covariance matrices to the deterministic Riccati flow is proven as the number of particles tends to infinity. Also presented are refined moment estimates and sharp bias and variance estimates. These expansions are also used to deduce a functional central limit theorem at the level of the diffusion process in matrix spaces.
We consider a class of finite-time horizon nonlinear stochastic optimal control problem. Although the optimal control admits a path integral representation for this class of control problems, efficient computation of the associated path integrals remains a challenging task. We propose a new Monte Carlo approach that significantly improves upon existing methodology. We tackle the issue of exponential growth in variance with the time horizon by casting optimal control estimation as a smoothing problem for a state-space model, and applying smoothing algorithms based on particle Markov chain Monte Carlo. To further reduce the cost, we then develop a multilevel Monte Carlo method which allows us to obtain an estimator of the optimal control with mean squared error with a cost of . In contrast, a cost of is required for the existing methodology to achieve the same mean squared error. Our approach is illustrated on two numerical examples.
The problem of distributed information (data) fusion is considered in this paper. Using a new class of functionals to perform data fusion – the power mean – information fusion can be performed in a dependency/correlation-agnostic manner while achieving better fusion results than many classical fusion techniques (i.e. linear or log-linear pooling). In addition, power mean techniques converge much more rapidly (even in finite time) in distributed network settings. Computing the power mean on generic probability distribution functions is specifically addressed in this paper, including a distributed protocol and specific steps for ensuring distributed convergence to the global power mean. Results demonstrate: 1). convergence to the global power mean even in a distributed setting, 2). the ability for certain power means to converge in finite time, and 3). the improved fusion results achieved by using the power mean over more traditional fusion techniques, even with relatively complex distributions.
The stability properties of matrix-valued Riccati diffusions are investigated. The matrix-valued Riccati diffusion processes considered in this work are of interest in their own right, as a rather prototypical model of a matrix-valued quadratic stochastic process. Under rather natural observability and controllability conditions, we derive time-uniform moment and fluctuation estimates and exponential contraction inequalities. Our approach combines spectral theory with nonlinear semigroup methods and stochastic matrix calculus. This analysis seem to be the first of its kind for this class of matrix-valued stochastic differential equation. This class of stochastic models arise in signal processing and data assimilation, and more particularly in ensemble Kalman-Bucy filtering theory. In this context, the Riccati diffusion represents the flow of the sample covariance matrices associated with McKean-Vlasov-type interacting Kalman-Bucy filters. The analysis developed here applies to filtering problems with unstable signals.
This article is concerned with the fluctuation analysis and the stability properties of a class of one-dimensional Riccati diffusions. These one-dimensional stochastic differential equations exhibit a quadratic drift function and a non-Lipschitz continuous diffusion function. We present a novel approach, combining tangent process techniques, Feynman-Kac path integration and exponential change of measures, to derive sharp exponential decays to equilibrium. We also provide uniform estimates with respect to the time horizon, quantifying with some precision the fluctuations of these diffusions around a limiting deterministic Riccati differential equation. These results provide a stronger and almost sure version of the conventional central limit theorem. We illustrate these results in the context of ensemble Kalman-Bucy filtering. To the best of our knowledge, the exponential stability and the fluctuation analysis developed in this work are the first results of this kind for this class of nonlinear diffusions.
This work is concerned with the stability properties of linear stochastic differential equations with random (drift and diffusion) coefficient matrices, and the stability of a corresponding random transition matrix (or exponential semigroup).We consider a class of random matrix drift coefficients that involves random perturbations of an exponentially stable flow of deterministic (time-varying) drift matrices.In contrast with more conventional studies, our analysis is not based on the existence of Lyapunov functions, and it does not rely on any ergodic properties.These approaches are often difficult to apply in practice when the drift/diffusion coefficients are random.We present rather weak and easily checked perturbation-type conditions for the asymptotic stability of time-varying and random linear stochastic differential equations.We provide new log-Lyapunov estimates and exponential contraction inequalities on any time horizon as soon as the fluctuation parameter is sufficiently small.These seem to be the first results of this type for this class of linear stochastic differential equations with random coefficient matrices.
These lecture notes provide a comprehensive, self-contained introduction to the analysis of Wishart matrix moments. This study may act as an introduction to some particular aspects of random matrix theory, or as a self-contained exposition of Wishart matrix moments. Random matrix theory plays a central role in statistical physics, computational mathematics and engineering sciences, including data assimilation, signal processing, combinatorial optimization, compressed sensing, econometrics and mathematical finance, among numerous others. The mathematical foundations of the theory of random matrices lies at the intersection of combinatorics, non-commutative algebra, geometry, multivariate functional and spectral analysis, and of course statistics and probability theory. As a result, most of the classical topics in random matrix theory are technical, and mathematically difficult to penetrate for non-experts and regular users and practitioners. The technical aim of these notes is to review and extend some important results in random matrix theory in the specific context of real random Wishart matrices. This special class of Gaussian-type sample covariance matrix plays an important role in multivariate analysis and in statistical theory. We derive non-asymptotic formulae for the full matrix moments of real valued Wishart random matrices. As a corollary, we derive and extend a number of spectral and trace-type results for the case of non-isotropic Wishart random matrices. We also derive the full matrix moment analogues of some classic spectral and trace-type moment results. For example, we derive semi-circle and Marchencko-Pastur-type laws in the non-isotropic and full matrix cases. Laplace matrix transforms and matrix moment estimates are also studied, along with new spectral and trace concentration-type inequalities.
This work considers a target tracking problem where the observed information is in the form of natural language-type statements. More specifically, the focus is on a spatio-temporal tracking problem where each uttered expression may involve both spatial, motion and temporal uncertainty, and a general modelling framework for natural language statements of a rather general semantic form is developed. This framework involves the definition of some tuple that allows one to extract the common semantics from arbitrary parsed expressions conveying some canonical information. Given this tuple, an estimation and tracking method based on the concept of outer probability measures is introduced and an estimation algorithm for handling this temporal uncertainty, along with delayed and out-of-sequence information arrival, is developed. This framework allows for modelling imprecise information in a more general and realistic sense.
Information or data fusion concerns the aggregation, or combination, of probability measures. For example, in machine learning, statistics and signal processing, one may seek to `combine' posterior distributions, [e.g. 1) Bayes classifiers or 2) posteriors over target states etc], arising from distinct but not necessarily independent sources. For example, sources might include partially disjoint trainers, or spatially distinct sensors correlated via state dependent measurements, etc. Data fusion is common in risk analysis where one is broadly interested in pooling expert opinions described by probability measures, and where it is often hard to assess and account for correlation among experts. The contribution of this work is the introduction of a broad class of data fusion rules that seek the combination of two (or more) probability distributions in the presence of non-zero, but unknown, correlation. We introduce rules that are improved in the sense that they are 'closer' to the true Bayesian result that would be computed if one could exploit knowledge of the correlation between the input distributions. We introduce these rules under the common algorithmic constraint of avoiding the so-called `double-counting' of correlated information. The general framework proposed is based on homogeneous functionals. We examine the fusion performance and computational properties when using these functionals. We also consider distributed data fusion on (possibly) time-varying and incomplete network topologies and related convergence properties.
This article is concerned with the fluctuation analysis and the stability properties of a class of one-dimensional Riccati diffusions. These one-dimensional stochastic differential equations exhibit a quadratic drift function and a non-Lipschitz continuous diffusion function. We present a novel approach, combining tangent process techniques, Feynman-Kac path integration, and exponential change of measures, to derive sharp exponential decays to equilibrium. We also provide uniform estimates with respect to the time horizon, quantifying with some precision the fluctuations of these diffusions around a limiting deterministic Riccati differential equation. These results provide a stronger and almost sure version of the conventional central limit theorem. We illustrate these results in the context of ensemble Kalman-Bucy filtering. To the best of our knowledge, the exponential stability and the fluctuation analysis developed in this work are the first results of this kind for this class of nonlinear diffusions.
This work considers the stability of nonlinear stochastic receding horizon control when the optimal controller is only computed approximately. A number of general classes of controller approximation error are analysed including deterministic and probabilistic errors and even controller sample and hold errors. In each case, it is shown that the controller approximation errors do not accumulate (even over an infinite time frame) and the process converges exponentially fast to a small neighbourhood of the origin. In addition to this analysis, an approximation method for receding horizon optimal control is proposed based on Monte Carlo simulation. This method is derived via the Feynman-Kac formula which gives a stochastic interpretation for the solution of a Hamilton-Jacobi-Bellman equation associated with the true optimal controller. It is shown, and it is a prime motivation for this study, that this particular controller approximation method practically stabilises the underlying nonlinear process.