In this note we consider the finite-dimensional parameter estimation problem associated to inverse problems. In such scenarios, one seeks to maximize the marginal likelihood associated to a Bayesian model. This latter model is connected to the solution of partial or ordinary differential equation. As such, there are two primary difficulties in maximizing the marginal likelihood (i) that the solution of differential equation is not always analytically tractable and (ii) neither is the marginal likelihood. Typically (i) is dealt with using a numerical solution of the differential equation, leading to a numerical bias and (ii) has been well studied in the literature using, for instance, Markovian stochastic approximation. It is well-known that to reduce the computational effort to obtain the maximal value of the parameter, one can use a hierarchy of solutions of the differential equation and combine with stochastic gradient methods. Several approaches do exactly this. In this paper we consider the asymptotic variance in the central limit theorem, associated to known estimates and find bounds on the asymptotic variance in terms of the precision of the solution of the differential equation. The significance of these bounds are the that they provide missing theoretical guidelines on how to set simulation parameters; that is, these appear to be the first mathematical results which help to run the methods efficiently in practice.
We consider the problem of Bayesian inference for bi-variate data observed in time but with observation times which occur non-synchronously. In particular, this occurs in a wide variety of applications in finance, such as high-frequency trading or crude oil futures trading. We adopt a diffusion model for the data and formulate a Bayesian model with priors on unknown parameters along with a latent representation for the the so-called missing data. We then consider computational methodology to fit the model using Markov chain Monte Carlo (MCMC). We have to resort to time-discretization methods as the complete data likelihood is intractable and this can cause considerable issues for MCMC when the data are observed in low frequencies. In a high frequency observation frequencies we present a simple particle MCMC method based on an Euler--Maruyama time discretization, which can be enhanced using multilevel Monte Carlo (MLMC). In the low frequency observation regime we introduce a novel bridging representation of the posterior in continuous time to deal with the issues of MCMC in this case. This representation is discretized and fitted using MCMC and MLMC. We apply our methodology to real and simulated data to establish the efficacy of our methodology.
We consider the discrete-time filtering problem in scenarios where the observation noise is degenerate or low. We focus on the case where the observation equation is a linear function of the state and that additive noise is low or degenerate, however, we place minimal assumptions on the hidden state process. In this scenario we derive new particle filtering (PF) algorithms and, under assumptions, in such a way that as the noise becomes more degenerate a PF which approximates the low noise filtering problem provably inherits the properties of the PF used in the degenerate case. We extend our framework to the case where the hidden states are drawn from a diffusion process. In this scenario we develop new PFs which are robust to both low noise and fine levels of time discretization. We illustrate our algorithms numerically on several examples.
In this paper we consider parameter estimation for discretely observed diffusion processes. In particular, we focus on data that are observed at low frequency and methodology that can estimate parameters with uncertainty quantification. Most statistical work in this domain develops advanced Markov chain Monte Carlo (MCMC) algorithms for sampling from the posterior of the parameters, a task which is often complicated by the fact that one seldom has access to the transition density of the diffusion process; one has to combine sophisticated MCMC methods which are robust to the required time discretization of the diffusion, which can yield expensive algorithms. We focus on developing the martingale posterior method for the context of interest, when one can only numerically approximate the transition density of the diffusion. Based on using types of diffusion bridges we introduce a new martingale posterior method for parameter estimation for discretely observed diffusion processes. We prove that this algorithm approximates, in some sense, the martingale posterior which has no time-discretization bias up-to 𝒪(Δ) if Δ is the time discretization step. Our approach is illustrated on several examples, showing orders of magnitude speed up versus state-of-the-art MCMC algorithms.
In this paper we consider the conditional stochastic optimization (CSO) problem. This consists of optimizing a function which can be written as the expectation of a function which is itself a function of a conditional expectation, i.e. of the type F(ξ) := 𝔼[f(Z,𝔼[g(Z,X,ξ)|Z])], where precise definitions are given in the main text. We address a particular class of CSO problems where the joint law of the random variables X,Z cannot be exactly sampled; this case has been addressed in Goda Kitade (2023). We introduce a method that combines Markovian stochastic approximation with unbiased approximation methods which allows one to find the optimizer of F(ξ) in the context of interest. We illustrate our methodology on two examples associated to parameter estimation with model averaging and portfolio selection associated to high-dimensional full factor multivariate stochastic volatility models.
In this paper we consider the parameter estimation problem associated to partially-observed time changed SDEs, with observations that are given at discrete times. In particular we consider both likelihood and Bayesian estimation. We develop new Markov chain Monte Carlo (MCMC) algorithms which allow an unbiased score-based stochastic approximation method to provide likelihood-type parameter estimators. We also use a variant of this MCMC algorithm to perform multilevel-based Bayesian parameter estimation. We prove that this latter method achieves a mean square error of 𝒪(ε^2) (ε>0) with a cost of 𝒪(ε^-2log(ε)^2). Our methodologies are tested numerically on both simulated and real data.
In this article we consider bridging between two mixture probability measures. In particular, given access to a Markov kernel between two component distributions, we provide a general mechanism to generate samples from one mixture to the other. Associated to a given reference and extended state space, we prove entropic optimality of this approach. In order to use this idea one needs to know the underlying mixtures and the Markov kernel, which is seldom available, and so we consider the case of Gaussian mixtures and Schrödinger Bridges. We prove a general 2-Wasserstein continuity bound between the exact bridge and one that is approximated, based on ε-covariance inflation, and these rely on a novel continuity analysis of perturbed Riccati maps. We apply our results in the context of bridging mixtures of Gaussians, single Gaussians and empirical estimators of the Gaussian parameters and the Monge map. For mixtures of Gaussians, when the parameters are estimated using the Expectation-Maxmization algorithm, the upper-bound on the 2-Wasserstein distance between the true and approximated bridges is, under assumptions and with probability at least 1-10N^-1, 𝒪([(dlog NN)^1/2{1+(dlog NN)^1/2(ε^-2+1)}+ε^2]) and for the other two cases, in expectation, 𝒪(d{1+ε^-21+N+ε^2}), where d,N∈ℕ is the dimension of the Gaussian and the number of empirical samples respectively. We also investigate our bounds numerically.
In this paper we investigate the efficacy of the score-based martingale posteriors (SMP) (Cui & Walker, 2025; Fong et al., 2023) in the context of modern and large-scale machine learning problems and its potential for meaningful uncertainty quantification. SMPs work with a stochastic gradient ascent-type recursion on the parameter space of stochastic models and construct a martingale on the parameter space. Under simple mathematical assumptions, the recursion can be built so that the parameters form a martingale sequence which possesses a limiting, in time, random variable, the latter of which can be simulated very quickly, in contrast to Monte Carlo-based methods such as Markov chain Monte Carlo. In this expository paper we explore the SMP for inferring the parameters of deep neural networks (DNNs) and, where feasible, compare our results to the state-of-the-art Monte Carlo methods aimed at inferring conventional Bayesian posteriors.
Entropic optimal transport problems play an increasingly important role in machine learning and generative modelling. In contrast with optimal transport maps which often have limited applicability in high dimensions, Schrödinger bridges can be solved using the celebrated Sinkhorn’s algorithm, a.k.a. the iterative proportional fitting procedure. The stability properties of Sinkhorn bridges when the number of iterations tends to infinity is a very active research area in applied probability and machine learning. Traditional proofs of convergence are mainly based on nonlinear versions of Perron–Frobenius theory and related Hilbert projective metric techniques, gradient descent, Bregman divergence techniques and Hamilton–Jacobi–Bellman equations, including propagation of convexity profiles based on coupling diffusions by reflection methods. The objective of this review article is to present, in a self-contained manner, recently developed Sinkhorn/Gibbs-type semigroup analysis based upon contraction coefficients and Lyapunov-type operator-theoretic techniques. These powerful, off-the-shelf semigroup methods are based upon transportation cost inequalities (e.g. log-Sobolev, Talagrand quadratic inequality, curvature estimates), ϕ -divergences, Kantorovich-type criteria and Dobrushin contraction-type coefficients on weighted Banach spaces as well as Wasserstein distances. This novel semigroup analysis allows one to unify and simplify many arguments in the stability of Sinkhorn’s algorithm. It also yields new contraction estimates with respect to generalized ϕ -entropies, as well as weighted total variation norms, Kantorovich criteria and Wasserstein distances.
In this paper we examine the numerical approximation of the limiting invariant measure associated with Feynman-Kac formulae. These are expressed in a discrete time formulation and are associated with a Markov chain and a potential function. The typical application considered here is the computation of eigenvalues associated with non-negative operators as found, for example, in physics or particle simulation of rare-events. We focus on a novel lagged approximation of this invariant measure, based upon the introduction of a ratio of time-averaged Feynman-Kac marginals associated with a positive operator iterated l is an element of N times; a lagged Feynman-Kac formula. This estimator and its approximation using Diffusion Monte Carlo (DMC) are commonly used in the physics literature. In short, DMC is an iterative algorithm involving N is an element of N particles or walkers simulated in parallel, that undergo sampling and resampling operations. In this work, it is shown that for the DMC approximation of the lagged Feynman-Kac formula, one has an almost sure characterization of the L1-error as the time parameter (iteration) goes to infinity and this is at most of O(exp{-xl}/N), for x > 0. In addition a non root asymptotic in time, and time uniform L1-bound is proved which is O(l/N). We also prove a novel central limit theorem to give a characterization of the exact asymptotic in time variance. This analysis demonstrates that the strategy used in physics, namely, to run DMC with N and l small and, for long time enough, is mathematically justified. Our results also suggest how one should choose N and l in practice. We emphasize that these results are not restricted to physical applications; they have broad relevance to the general problem of particle simulation of the Feynman-Kac formula, which is utilized in a great variety of scientific and engineering fields.
In this paper we consider the estimation of unknown parameters in Bayesian inverse problems. In most cases of practical interest, there are several barriers to performing such estimation, This includes a numerical approximation of a solution of a differential equation and, even if exact solutions are available, an analytical intractability of the marginal likelihood and its associated gradient, which is used for parameter estimation. The focus of this article is to deliver unbiased estimates of the unknown parameters, that is, stochastic estimators that, in expectation, are equal to the maximize of the marginal likelihood, and possess no numerical approximation error. Based upon the ideas of [4] we develop a new approach for unbiased parameter estimation for Bayesian inverse problems. We prove unbiasedness and establish numerically that the associated estimation procedure is faster than the current state-of-the-art methodology for this problem. We demonstrate the performance of our methodology on a range of problems which include a PDE and ODE.
In this article we consider the estimation of static parameters for partially observed diffusion processes with discrete-time observations over a fixed time interval. In particular, when one only has access to time-discretized solutions of the diffusions we build upon the works of to devise a method that can estimate the parameters without time-discretization bias. We leverage an identity associated to the gradient of the log-likelihood associated to diffusion bridges, which has not been used before. Contrary to the afore mentioned methods, the diffusion coefficient can depend on the parameters and our approach facilitates the use of more efficient Markov chain sampling algorithms. We prove that our estimator is unbiased with finite variance and demonstrate the efficacy of our methodology in several examples.
Abstract. In this article we consider Bayesian inference associated to deep neural networks (DNNs) and in particular, trace-class neural network (TNN) priors [T. Sell and S. S. Singh, Dimension-robust Function Space Mcmc with Neural Network Priors, 2020] which can be preferable to traditional DNNs because they (a) are identifiable and (b) possess desirable convergence properties. TNN priors are defined on functions with infinitely many hidden units, and have strongly convergent Karhunen–Loeve-type approximations with finitely many hidden units. A practical hurdle is that the Bayesian solution is computationally demanding, requiring simulation methods, so approaches to drive down the complexity are needed. In this paper, we leverage the strong convergence of TNN in order to apply multilevel Monte Carlo (MLMC) to these models. In particular, an MLMC method that was introduced by Beskos et al. [ SIAM/ASA J. Uncertain. Quantif., 6 (2018), pp. 762–786] is used to approximate posterior expectations of Bayesian TNN models with optimal computational complexity, and this is mathematically proved. The results are verified with several numerical experiments on model problems arising in machine learning, including some toy regression and classification models, MNIST image classification, and a challenging reinforcement learning problem. Furthermore, we illustrate the practical utility of the method on MNIST as well as IMDb sentiment classification.
In this article, we consider the modeling of measurement error for fund returns data. In particular, given access to a time series of discretely observed log-returns and the associated maximum over the observation period, we develop a stochastic model that models the true log-returns and maximum via a L & eacute;vy process and the data as a measurement error thereof. The main technical difficulty of trying to infer this model, for instance Bayesian parameter estimation, is that the joint transition density of the return and maximum is seldom known, nor can it be simulated exactly. Based upon the novel stick-breaking representation of [1 c1], we provide an approximation of the model. We develop a Markov chain Monte Carlo (MCMC) algorithm to sample from the Bayesian posterior of the approximated posterior and then extend this to a multilevel MCMC method, which can reduce the computational cost to approximate posterior expectations, relative to ordinary MCMC. We prove that the computational complexity of our multilevel MCMC scheme is optimal for L & eacute;vy models with increments that can be sampled in constant time. We implement our methodology on several applications, including for real data.
We present a new antithetic multilevel Monte Carlo (MLMC) method for the estimation of expectations with respect to laws of diffusion processes that can be elliptic or hypo elliptic. In particular, we consider the case where one has to resort to time discretization of the diffusion and numerical simulation of such schemes. Inspired by recent works, we introduce a new MLMC estimator of expectations, which does not require any Le'\vy area simulation and has a strong error of order 2 and a weak error of order 2. We then show how this approach can be used in the context of the filtering problem associated with partially observed diffusions with discrete time observations. We illustrate that in numerical simulations our new approaches provide efficiency gains for several problems, particularly when the diffusion process is hypo elliptic, relative to some existing methods.
In this article we consider Bayesian estimation of static parameters for a class of partially observed McKean-Vlasov diffusion processes with discrete-time observations over a fixed time interval. This problem features several obstacles to its solution, which include that the posterior density is numerically intractable in continuous-time, even if the transition probabilities are available and even when one uses a time-discretization, the posterior still cannot be used by adopting well-known computational methods such as Markov chain Monte Carlo (MCMC). In this paper we provide a solution to this problem by using new MCMC algorithms which can solve the afore-mentioned issues. This MCMC algorithm is extended to use multilevel Monte Carlo (MLMC) methods. We prove convergence bounds on our parameter estimators and show that the MLMC-based MCMC algorithm reduces the computational cost to achieve a mean square error versus ordinary MCMC by an order of magnitude. We numerically illustrate our results on two models.
In this article we consider Bayesian inference associated to deep neural networks (DNNs) and in particular, trace-class neural network (TNN) priors which can be preferable to traditional DNNs as (a) they are identifiable and (b) they possess desirable convergence properties. TNN priors are defined on functions with infinitely many hidden units, and have strongly convergent Karhunen-Loeve-type approximations with finitely many hidden units. A practical hurdle is that the Bayesian solution is computationally demanding, requiring simulation methods, so approaches to drive down the complexity are needed. In this paper, we leverage the strong convergence of TNN in order to apply Multilevel Monte Carlo (MLMC) to these models. In particular, an MLMC method that was introduced is used to approximate posterior expectations of Bayesian TNN models with optimal computational complexity, and this is mathematically proved. The results are verified with several numerical experiments on model problems arising in machine learning, including regression, classification, and reinforcement learning.
In this paper we consider Bayesian parameter inference associated to a class of partially observed stochastic differential equations (SDE) driven by jump processes. Such type of models can be routinely found in applications, of which we focus upon the case of neuroscience. The data are assumed to be observed regularly in time and driven by the SDE model with unknown parameters. In practice the SDE may not have an analytically tractable solution and this leads naturally to a time-discretization. We adapt the multilevel Markov chain Monte Carlo method of [11], which works with a hierarchy of time discretizations and show empirically and theoretically that this is preferable to using one single time discretization. The improvement is in terms of the computational cost needed to obtain a pre-specified numerical error. Our approach is illustrated on models that are found in neuroscience.
In this article, we consider the problem of clustering multi-view data, that is, information associated to individuals that form heterogeneous data sources (the views). We adopt a Bayesian model and in the prior structure we assume that each individual belongs to a baseline cluster and conditionally allow each individual in each view to potentially belong to different clusters than the baseline. We call such a structure ''latent modularity''. Then for each cluster, in each view we have a specific statistical model with an associated prior. We derive expressions for the marginal priors on the view-specific cluster labels and the associated partitions, giving several insights into our chosen prior structure. Using simple Markov chain Monte Carlo algorithms, we consider our model in a simulation study, along with a more detailed case study that requires several modeling innovations.
We consider the long time behavior of Wong-Zakai approximations of stochastic differential equations. These piecewise smooth diffusion approximations are of great importance in many areas, such as those with ordinary differential equations associated to random smooth fluctuations; e.g. robust filtering problems. In many examples, the mean error estimate bounds that have been derived in the literature can grow exponentially with respect to the time horizon. We show in a simple example that indeed mean error estimates do explode exponentially in the time parameter, i.e. in that case a Wong-Zakai approximation is only useful for extremely short time intervals. Under spectral conditions, we present some quantitative time-uniform convergence theorems, i.e. time-uniform mean error bounds, yielding what seems to be the first results of this type for Wong-Zakai diffusion approximations.