In recent years, Bayesian inference in large-scale inverse problems found in science, engineering and machine learning has gained significant attention. This paper examines the robustness of the Bayesian approach by analyzing the stability of posterior measures in relation to perturbations in the likelihood potential and the prior measure. We present new stability results using a family of integral probability metrics (divergences) akin to dual problems that arise in optimal transport. Our results stand out from previous works in three directions: (1) We construct new families of integral probability metrics that are adapted to the problem at hand; (2) These new metrics allow us to study both likelihood and prior perturbations in a convenient way; and (3) our analysis accommodates likelihood potentials that are only locally Lipschitz, making them applicable to a wide range of nonlinear inverse problems. Our theoretical findings are further reinforced through specific and novel examples where the approximation rates of posterior measures are obtained for different types of perturbations and provide a path towards the convergence analysis of recently adapted machine learning techniques for Bayesian inverse problems such as data-driven priors and neural network surrogates.
Estimating the statistics of the state of a dynamical system, from partial and noisy observations, is both mathematically challenging and finds wide application. Furthermore, the applications are of great societal importance, including problems such as probabilistic weather forecasting and prediction of epidemics. Particle filters provide a well-founded approach to the problem, leading to provably accurate approximations of the statistics. However these methods perform poorly in high dimensions. In 1994 the idea of ensemble Kalman filtering was introduced by Evensen, leading to a methodology that has been widely adopted in the geophysical sciences and also finds application to quite general inverse problems. However, ensemble Kalman filters have defied rigorous analysis of their statistical accuracy, except in the linear Gaussian setting. In this article we describe recent work which takes first steps to analyze the statistical accuracy of ensemble Kalman filters beyond the linear Gaussian setting. The subject is inherently technical, as it involves the evolution of probability measures according to a nonlinear and nonautonomous dynamical system; and the approximation of this evolution. It can nonetheless be presented in a fairly accessible fashion, understandable with basic knowledge of dynamical systems, numerical analysis and probability.
For algorithms based on interacting particle systems that admit a mean-field description, convergence analysis is often more accessible at the mean-field level. In order to transfer convergence results obtained at the mean-field level to the finite ensemble size setting, it is desirable to show that the particle dynamics converge in an appropriate sense to the corresponding mean-field dynamics. In this paper, we prove quantitative mean-field limit results for two related interacting particle systems: Consensus-Based Optimization and Consensus-Based Sampling. Our approach requires a generalization of Sznitman's classical argument: in order to circumvent issues related to the lack of global Lipschitz continuity of the coefficients, we discard an event of small probability, the contribution of which is controlled using moment estimates for the particle systems. In addition, we present new results on the well-posedness of the particle systems and their mean-field limit, and provide novel stability estimates for the weighted mean and the weighted covariance.
The workshop brought together researchers from geometry, nonlinear functional analysis, calculus of variations, partial differential equations, and stochastics around a common topic: systems whose evolution is driven by variational principles such as gradient or Hamiltonian systems. The talks covered a wide range of topics, including variational tools such as incremental minimization approximations, Gamma convergence, and optimal transport, reaction-diffusion systems, singular perturbation and homogenization, rate-independent models for visco-plasticity and fracture, Hamiltonian and hyperbolic systems, stochastic models and new gradient structures for Markov processes or variational large-deviation principles.
The ensemble Kalman filter is widely used in applications because, for high dimensional filtering problems, it has a robustness that is not shared for example by the particle filter; in particular it does not suffer from weight collapse. However, there is no theory which quantifies its accuracy as an approximation of the true filtering distribution, except in the Gaussian setting. To address this issue we provide the first analysis of the accuracy of the ensemble Kalman filter beyond the Gaussian setting. We prove two types of results: the first type comprise a stability estimate controlling the error made by the ensemble Kalman filter in terms of the difference between the true filtering distribution and a nearby Gaussian; and the second type use this stability result to show that, in a neighbourhood of Gaussian problems, the ensemble Kalman filter makes a small error, in comparison with the true filtering distribution. Our analysis is developed for the mean field ensemble Kalman filter. We rewrite the update equations for this filter, and for the true filtering distribution, in terms of maps on probability measures. We introduce a weighted total variation metric to estimate the distance between the two filters and we prove various stability estimates for the maps defining the evolution of the two filters, in this metric. Using these stability estimates we prove results of the first and second types, in the weighted total variation metric. We also provide a generalization of these results to the Gaussian projected filter, which can be viewed as a mean field description of the unscented Kalman filter.
We perform stability analysis of a kinetic bacterial chemotaxis model of bacterial self-organization, assuming that bacteria respond sharply to chemical signals. The resulting discontinuous tumbling kernel represents the key challenge for the stability analysis as it rules out a direct linearization of the nonlinear terms. To address this challenge we fruitfully separate the evolution of the shape of the cellular profile from its global motion. We provide a full nonlinear stability theorem in a perturbative setting when chemical degradation can be neglected. With chemical degradation we prove stability of the linearized operator. In both cases we obtain exponential relaxation to equilibrium with an explicit rate using hypocoercivity techniques. To apply a hypocoercivity approach in this setting, we develop two novel and specific approaches: i) the use of the H^1 norm instead of the L^2 norm, and ii) the treatment of nonlinear terms. This work represents an important step forward in bacterial chemotaxis modeling from a kinetic perspective as most results are currently only available for the macroscopic descriptions, which are usually parabolic in nature. Significant difficulty arises due to the lack of regularization of the kinetic transport operator as compared to the parabolic operator in the macroscopic scaling limit.
We study a variant of the dynamical optimal transport problem in which the energy to be minimised is modulated by the covariance matrix of the distribution. Such transport metrics arise naturally in mean-field limits of certain ensemble Kalman methods for solving inverse problems. We show that the transport problem splits into two coupled minimization problems: one for the evolution of mean and covariance of the interpolating curve and one for its shape. The latter consists in minimising the usual Wasserstein length under the constraint of maintaining fixed mean and covariance along the interpolation. We analyse the geometry induced by this modulated transport distance on the space of probabilities as well as the dynamics of the associated gradient flows. Those show better convergence properties in comparison to the classical Wasserstein metric in terms of exponential convergence rates independent of the Gaussian target. On the level of the gradient flows a similar splitting into the evolution of moments and shapes of the distribution can be observed.
We perform the nonlinear stability analysis of a chemotaxis model of bacterial self-organization, assuming that bacteria respond sharply to chemical signals. The resulting discontinuous advection speed represents the key challenge for the stability analysis. We follow a perturbative approach, where the shape of the cellular profile is clearly separated from its global motion, allowing us to circumvent the discontinuity issue. Further, the homogeneity of the problem leads to two conservation laws, which express themselves in differently weighted functional spaces. This discrepancy between the weights represents another key methodological challenge. We derive an improved Poincaré inequality that allows to transfer the information encoded in the conservation laws to the appropriately weighted spaces. As a result, we obtain exponential relaxation to equilibrium with an explicit rate. A numerical investigation illustrates our results.
We present a model-free data-driven inference method that enables inferences on system outcomes to be derived directly from empirical data without the need for intervening modeling of any type, be it modeling of a material law or modeling of a prior distribution of material states. We specifically consider physical systems with states characterized by points in a phase space determined by the governing field equations. We assume that the system is characterized by two likelihood measures: one $$\mu _D$$ measuring the likelihood of observing a material state in phase space; and another $$\mu _E$$ measuring the likelihood of states satisfying the field equations, possibly under random actuation. We introduce a notion of intersection between measures which can be interpreted to quantify the likelihood of system outcomes. We provide conditions under which the intersection can be characterized as the athermal limit $$\mu _\infty $$ of entropic regularizations $$\mu _\beta $$ , or thermalizations, of the product measure $$\mu = \mu _D\times \mu _E$$ as $$\beta \rightarrow +\infty $$ . We also supply conditions under which $$\mu _\infty $$ can be obtained as the athermal limit of carefully thermalized $$(\mu _{h,\beta _h})$$ sequences of empirical data sets $$(\mu _h)$$ approximating weakly an unknown likelihood function $$\mu $$ . In particular, we find that the cooling sequence $$\beta _h \rightarrow +\infty $$ must be slow enough, corresponding to annealing, in order for the proper limit $$\mu _\infty $$ to be delivered. Finally, we derive explicit analytic expressions for expectations $$\mathbb {E}[f]$$ of outcomes f that are explicit in the data, thus demonstrating the feasibility of the model-free data-driven paradigm as regards making convergent inferences directly from the data without recourse to intermediate modeling steps.
We propose a novel framework for analyzing the dynamics of distribution shift in real-world systems that captures the feedback loop between learning algorithms and the distributions on which they are deployed. Prior work largely models feedback-induced distribution shift as adversarial or via an overly simplistic distribution-shift structure. In contrast, we propose a coupled partial differential equation model that captures fine-grained changes in the distribution over time by accounting for complex dynamics that arise due to strategic responses to algorithmic decision-making, non-local endogenous population interactions, and other exogenous sources of distribution shift. We consider two common settings in machine learning: cooperative settings with information asymmetries, and competitive settings where a learner faces strategic users. For both of these settings, when the algorithm retrains via gradient descent, we prove asymptotic convergence of the retraining procedure to a steady-state, both in finite and in infinite dimensions, obtaining explicit rates in terms of the model parameters. To do so we derive new results on the convergence of coupled PDEs that extends what is known on multi-species systems. Empirically, we show that our approach captures well-documented forms of distribution shifts like polarization and disparate impacts that simpler models cannot capture.
We study a Data-Driven approach to inference in physical systems in a measure-theoretic framework. The systems under consideration are characterized by two measures defined over the phase space: (i) A physical likelihood measure expressing the likelihood that a state of the system be admissible, in the sense of satisfying all governing physical laws; (ii) A material likelihood measure expressing the likelihood that a local state of the material be observed in the laboratory. We assume deterministic loading, which means that the first measure is supported on a linear subspace. We additionally assume that the second measure is only known approximately through a sequence of empirical (discrete) measures. We develop a method for the quantitative analysis of convergence based on the flat metric and obtain error bounds both for annealing and the discretization or sampling procedure, leading to the determination of appropriate quantitative annealing rates. Finally, we provide an example illustrating the application of the theory to transportation networks.
Graph Laplacians computed from weighted adjacency matrices are widely used to identify geometric structure in data, and clusters in particular; their spectral properties play a central role in a number of unsupervised and semi-supervised learning algorithms. When suitably scaled, graph Laplacians approach limiting continuum operators in the large data limit. Studying these limiting operators, therefore, sheds light on learning algorithms. This paper is devoted to the study of a parameterized family of divergence form elliptic operators that arise as the large data limit of graph Laplacians. The link between a three-parameter family of graph Laplacians and a three-parameter family of differential operators is explained. The spectral properties of these differential operators are analyzed in the situation where the data comprises two nearly separated clusters, in a sense which is made precise. In particular, we investigate how the spectral gap depends on the three parameters entering the graph Laplacian, and on a parameter measuring the size of the perturbation from the perfectly clustered case. Numerical results are presented which exemplify and extend the analysis: the computations study situations in which there are two nearly separated clusters, but which violate the assumptions used in our theory; situations in which more than two clusters are present, also going beyond our theory; and situations which demonstrate the relevance of our studies of differential operators for the understanding of finite data problems via the graph Laplacian. The findings provide insight into parameter choices made in learning algorithms which are based on weighted adjacency matrices; they also provide the basis for analysis of the consistency of various unsupervised and semi-supervised learning algorithms, in the large data limit.
The ensemble Kalman filter is widely used in applications because, for high dimensional filtering problems, it has a robustness that is not shared for example by the particle filter; in particular it does not suffer from weight collapse. However, there is no theory which quantifies its accuracy as an approximation of the true filtering distribution, except in the Gaussian setting. To address this issue we provide the first analysis of the accuracy of the ensemble Kalman filter beyond the Gaussian setting. Our analysis is developed for the mean field ensemble Kalman filter. We rewrite this filter in terms of maps on probability measures, and then we prove that these maps are locally Lipschitz in an appropriate weighted total variation metric. Using these stability estimates we demonstrate that, if the true filtering distribution is close to Gaussian after appropriate lifting to the joint space of state and data, then it is well approximated by the ensemble Kalman filter. Finally, we provide a generalization of these results to the Gaussian projected filter, which can be viewed as a mean field description of the unscented Kalman filter.
We propose a novel method for sampling and optimization tasks based on a stochastic interacting particle system. We explain how this method can be used for the following two goals: (i) generating approximate samples from a given target distribution; (ii) optimizing a given objective function. The approach is derivative-free and affine invariant, and is therefore well-suited for solving inverse problems defined by complex forward models: (i) allows generation of samples from the Bayesian posterior and (ii) allows determination of the maximum a posteriori estimator. We investigate the properties of the proposed family of methods in terms of various parameter choices, both analytically and by means of numerical simulations. The analysis and numerical simulation establish that the method has potential for general purpose optimization tasks over Euclidean space; contraction properties of the algorithm are established under suitable conditions, and computational experiments demonstrate wide basins of attraction for various specific problems. The analysis and experiments also demonstrate the potential for the sampling methodology in regimes in which the target distribution is unimodal and close to Gaussian; indeed we prove that the method recovers a Laplace approximation to the measure in certain parametric regimes and provide numerical evidence that this Laplace approximation attracts a large set of initial conditions in a number of examples.
Knowledge Graph Embeddings (KGEs) have shown promising performance on link prediction tasks by mapping the entities and relations from a knowledge graph into a geometric space. The capability of KGEs in preserving graph characteristics including structural aspects and semantics, highly depends on the design of their score function, as well as the inherited abilities from the underlying geometry. Many KGEs use the Euclidean geometry which renders them incapable of preserving complex structures and consequently causes wrong inferences by the models. To address this problem, we propose a neuro differential KGE that embeds nodes of a KG on the trajectories of Ordinary Differential Equations (ODEs). To this end, we represent each relation (edge) in a KG as a vector field on several manifolds. We specifically parameterize ODEs by a neural network to represent complex manifolds and complex vector fields on the manifolds. Therefore, the underlying embedding space is capable to assume the shape of various geometric forms to encode heterogeneous subgraphs. Experiments on synthetic and benchmark datasets using state-of-the-art KGE models justify the ODE trajectories as a means to enable structure preservation and consequently avoiding wrong inferences.
We present a model-free data-driven inference method that enables inferences on system outcomes to be derived directly from empirical data without the need for intervening modeling of any type, be it modeling of a material law or modeling of a prior distribution of material states. We specifically consider physical systems with states characterized by points in a phase space determined by the governing field equations. We assume that the system is characterized by two likelihood measures: one $\mu_D$ measuring the likelihood of observing a material state in phase space; and another $\mu_E$ measuring the likelihood of states satisfying the field equations, possibly under random actuation. We introduce a notion of intersection between measures which can be interpreted to quantify the likelihood of system outcomes. We provide conditions under which the intersection can be characterized as the athermal limit $\mu_\infty$ of entropic regularizations $\mu_\beta$, or thermalizations, of the product measure $\mu = \mu_D\times \mu_E$ as $\beta \to +\infty$. We also supply conditions under which $\mu_\infty$ can be obtained as the athermal limit of carefully thermalized $(\mu_{h,\beta_h})$ sequences of empirical data sets $(\mu_h)$ approximating weakly an unknown likelihood function $\mu$. In particular, we find that the cooling sequence $\beta_h \to +\infty$ must be slow enough, corresponding to quenching, in order for the proper limit $\mu_\infty$ to be delivered. Finally, we derive explicit analytic expressions for expectations $\mathbb{E}[f]$ of outcomes $f$ that are explicit in the data, thus demonstrating the feasibility of the model-free data-driven paradigm as regards making convergent inferences directly from the data without recourse to intermediate modeling steps.
We analyze the spectral clustering procedure for identifying coarse structure in a data set x₁,…,x_n, and in particular study the geometry of graph Laplacian embeddings which form the basis for spectral clustering algorithms. More precisely, we assume that the data is sampled from a mixture model supported on a manifold M embedded in R^d, and pick a connectivity length-scale e>0 to construct a kernelized graph Laplacian. We introduce a notion of a well-separated mixture model which only depends on the model itself, and prove that when the model is well separated, with high probability the embedded data set concentrates on cones that are centered around orthogonal vectors. Our results are meaningful in the regime where e=e(n) is allowed to decay to zero at a slow enough rate as the number of data points grows. This rate depends on the intrinsic dimension of the manifold on which the data is supported.
We consider a generalised Keller-Segel model with non-linear porous medium type diffusion and non-local attractive power law interaction, focusing on potentials that are more singular than Newtonian interaction. We show uniqueness of stationary states (if they exist) in any dimension both in the diffusion-dominated regime and in the fair-competition regime when attraction and repulsion are in balance. As stationary states are radially symmetric decreasing, the question of uniqueness reduces to the radial setting. Our key result is a sharp generalised Hardy-Littlewood-Sobolev type functional inequality in the radial setting.
Graph-based semi-supervised learning is the problem of propagating labels from a small number of labelled data points to a larger set of unlabelled data. This paper is concerned with the consistency of optimization-based techniques for such problems, in the limit where the labels have small noise and the underlying unlabelled data is well clustered. We study graph-based probit for binary classification, and a natural generalization of this method to multi-class classification using one-hot encoding. The resulting objective function to be optimized comprises the sum of a quadratic form defined through a rational function of the graph Laplacian, involving only the unlabelled data, and a fidelity term involving only the labelled data. The consistency analysis sheds light on the choice of the rational function defining the optimization.
Solving inverse problems without the use of derivatives or adjoints of the forward model is highly desirable in many applications arising in science and engineering. In this paper, we propose a new version of such a methodology, a framework for its analysis, and numerical evidence of the practicality of the method proposed. Our starting point is an ensemble of over-damped Langevin diffusions which interact through a single preconditioner computed as the empirical ensemble covariance. We demonstrate that the nonlinear Fokker-Planck equation arising from the mean-field limit of the associated stochastic differential equation (SDE) has a novel gradient flow structure, built on the Wasserstein metric and the covariance matrix of the noisy flow. Using this structure, we investigate large time properties of the Fokker-Planck equation, showing that its invariant measure coincides with that of a single Langevin diffusion, and demonstrating exponential convergence to the invariant measure in a number of settings. We introduce a new noisy variant on ensemble Kalman inversion (EKI) algorithms found from the original SDE by replacing exact gradients with ensemble differences; this defines the ensemble Kalman sampler (EKS). Numerical results are presented which demonstrate its efficacy as a derivative-free approximate sampler for the Bayesian posterior arising from inverse problems.