Using i.i.d. data to estimate a high-dimensional distribution in Wasserstein distance is a fundamental instance of the curse of dimensionality. We explore how structural knowledge about the data-generating process which gives rise to the distribution can be used to overcome this curse. More precisely, we work with the set of distributions of probabilistic graphical models for a known directed acyclic graph. It turns out that this knowledge is only helpful if it can be quantified, which we formalize via smoothness conditions on the transition kernels in the disintegration corresponding to the graph. In this case, we prove that the rate of estimation is governed by the local structure of the graph, more precisely by dimensions corresponding to single nodes together with their parent nodes. The precise rate depends on the exact notion of smoothness assumed for the kernels, where either weak (Wasserstein-Lipschitz) or strong (bidirectional Total-Variation-Lipschitz) conditions lead to different results. We prove sharpness under the strong condition and show that this condition is satisfied for example for distributions having a positive Lipschitz density.
We study stability and sample complexity properties of divergence regularized optimal transport (DOT). First, we obtain quantitative stability results for optimizers of DOT measured in Wasserstein distance, which are applicable to a wide class of divergences and simultaneously improve known results for entropic optimal transport. Second, we study the case of sample complexity, where the DOT problem is approximated using empirical measures of the marginals. We show that divergence regularization can improve the corresponding convergence rate compared to unregularized optimal transport. To this end, we prove upper bounds which exploit both the regularity of cost function and divergence functional, as well as the intrinsic dimension of the marginals. Along the way, we establish regularity properties of dual optimizers of DOT, as well as general limit theorems for empirical measures with suitable classes of test functions.
Entropic optimal transport – the optimal transport problem regularized by KL divergence – is highly successful in statistical applications. Thanks to the smoothness of the entropic coupling, its sample complexity avoids the curse of dimensionality suffered by unregularized optimal transport. The flip side of smoothness is overspreading: the entropic coupling always has full support, whereas the unregularized coupling that it approximates is usually sparse, even given by a map. Regularizing optimal transport by less-smooth f-divergences such as Tsallis divergence (i.e., L^p-regularization) is known to allow for sparse approximations, but is often thought to suffer from the curse of dimensionality as the couplings have limited differentiability and the dual is not strongly concave. We refute this conventional wisdom and show, for a broad family of divergences, that the key empirical quantities converge at the parametric rate, independently of the dimension. More precisely, we provide central limit theorems for the optimal cost, the optimal coupling, and the dual potentials induced by i.i.d. samples from the marginals. These results are obtained by a powerful yet elementary approach that is of broader interest for Z-estimation in function classes that are not Donsker.
Motivated by the entropic optimal transport problem in unbounded settings, we study versions of Hilbert's projective metric for spaces of integrable functions of bounded growth. These versions of Hilbert's metric originate from cones which are relaxations of the cone of all non-negative functions, in the sense that they include all functions having non-negative integral values when multiplied with certain test functions. We show that kernel integral operators are contractions with respect to suitable specifications of such metrics even for kernels which are not bounded away from zero, provided that the decay to zero of the kernel is controlled. As an application to entropic optimal transport, we show exponential convergence of Sinkhorn's algorithm in settings where the marginal distributions have sufficiently light tails compared to the growth of the cost function.
In this paper, we introduce a variant of optimal transport adapted to the causal structure given by an underlying directed graph $G$. Different graph structures lead to different specifications of the optimal transport problem. For instance, a fully connected graph yields standard optimal transport, a linear graph structure corresponds to causal optimal transport between the distributions of two discrete-time stochastic processes, and an empty graph leads to a notion of optimal transport related to CO-OT, Gromov-Wasserstein distances and factored OT. We derive different characterizations of $G$-causal transport plans and introduce Wasserstein distances between causal models that respect the underlying graph structure. We show that average treatment effects are continuous with respect to $G$-causal Wasserstein distances and small perturbations of structural causal models lead to small deviations in $G$-causal Wasserstein distance. We also introduce an interpolation between causal models based on $G$-causal Wasserstein distance and compare it to standard Wasserstein interpolation.
We address the problem of estimating the expected shortfall risk of a financial loss using a finite number of i.i.d. data. It is well known that the classical plug-in estimator suffers from poor statistical performance when faced with (heavy-tailed) distributions that are commonly used in financial contexts. Further, it lacks robustness, as the modification of even a single data point can cause a significant distortion. We propose a novel procedure for the estimation of the expected shortfall and prove that it recovers the best possible statistical properties (dictated by the central limit theorem) under minimal assumptions and for all finite numbers of data. Further, this estimator is adversarially robust: even if a (small) proportion of the data is maliciously modified, the procedure continuous to optimally estimate the true expected shortfall risk. We demonstrate that our estimator outperforms the classical plug-in estimator through a variety of numerical experiments across a range of standard loss distributions.
We study the convergence of divergence-regularized optimal transport as the regularization parameter vanishes. Sharp rates for general divergences including relative entropy or L p regularization, general transport costs, and multimarginal problems are obtained. A novel methodology using quantization and martingale couplings is suitable for noncompact marginals and achieves, in particular, the sharp leading-order term of entropically regularized 2-Wasserstein distance for marginals with a finite [Formula: see text]-moment. Funding: This work was supported by the Alfred P. Sloan Foundation [Grant FG-2016-6282] and the Division of Mathematical Sciences [Grants DMS-1812661 and DMS-2106056].
Adapted optimal transport (AOT) problems are optimal transport problems for distributions of a time series where couplings are constrained to have a temporal causal structure. In this paper, we develop computational tools for solving AOT problems numerically. First, we show that AOT problems are stable with respect to perturbations in the marginals and thus arbitrary AOT problems can be approximated by sequences of linear programs. We further study entropic methods to solve AOT problems. We show that any entropically regularized AOT problem converges to the corresponding unregularized problem if the regularization parameter goes to zero. The proof is based on a novel method - even in the non-adapted case - to easily obtain smooth approximations of a given coupling with fixed marginals. Finally, we show tractability of the adapted version of Sinkhorn's algorithm. We give explicit solutions for the occurring projections and prove that the procedure converges to the optimizer of the entropic AOT problem.
We study versions of Hilbert's projective metric for spaces of integrable functions of bounded growth. These metrics originate from cones which are relaxations of the cone of all non-negative functions, in the sense that they include all functions having non-negative integral values when multiplied with certain test functions. We show that kernel integral operators are contractions with respect to suitable specifications of such metrics even for kernels which are not bounded away from zero, provided that the decay to zero of the kernel is controlled. As an application to entropic optimal transport, we show exponential convergence of Sinkhorn's algorithm in settings where the marginal distributions have sufficiently light tails compared to the growth of the cost function.
In a high-dimensional regression framework, we study consequences of the naive two-step procedure where first the dimension of the input variables is reduced and second, the reduced input variables are used to predict the output variable with kernel regression. In order to analyze the resulting regression errors, a novel stability result for kernel regression with respect to the Wasserstein distance is derived. This allows us to bound errors that occur when perturbed input data is used to fit the regression function. We apply the general stability result to principal component analysis (PCA). Exploiting known estimates from the literature on both principal component analysis and kernel regression, we deduce convergence rates for the two-step procedure. The latter turns out to be particularly useful in a semi-supervised setting.
In the theory of lossy compression, the rate-distortion (R-D) function $R(D)$ describes how much a data source can be compressed (in bit-rate) at any given level of fidelity (distortion). Obtaining $R(D)$ for a given data source establishes the fundamental performance limit for all compression algorithms. We propose a new method to estimate $R(D)$ from the perspective of optimal transport. Unlike the classic Blahut--Arimoto algorithm which fixes the support of the reproduction distribution in advance, our Wasserstein gradient descent algorithm learns the support of the optimal reproduction distribution by moving particles. We prove its local convergence and analyze the sample complexity of our R-D estimator based on a connection to entropic optimal transport. Experimentally, we obtain comparable or tighter bounds than state-of-the-art neural network methods on low-rate sources while requiring considerably less tuning and computation effort. We also highlight a connection to maximum-likelihood deconvolution and introduce a new class of sources that can be used as test cases with known solutions to the R-D problem.
We study the stability of entropically regularized optimal transport with respect to the marginals. Lipschitz continuity of the value and Hölder continuity of the optimal coupling in p-Wasserstein distance are obtained under general conditions including quadratic costs and unbounded marginals. The results for the value extend to regularization by an arbitrary divergence. As an application, we show convergence of Sinkhorn's algorithm in Wasserstein sense, including for quadratic cost. Two techniques are presented: The first compares an optimal coupling with its so-called shadow, a coupling induced on other marginals by an explicit construction. The second transforms one set of marginals by a change of coordinates and thus reduces the comparison of differing marginals to the comparison of differing cost functions under the same marginals.
Motivated by applications in model-free finance and quantitative risk management, we consider Fr\'echet classes of multivariate distribution functions where additional information on the joint distribution is assumed, while uncertainty in the marginals is also possible. We derive optimal transport duality results for these Fr\'echet classes that extend previous results in the related literature. These proofs are based on representation results for increasing convex functionals and the explicit computation of the conjugates. We show that the dual transport problem admits an explicit solution for the function $f=1_B$, where $B$ is a rectangular subset of $\mathbb R^d$, and provide an intuitive geometric interpretation of this result. The improved Fr\'echet--Hoeffding bounds provide ad-hoc upper bounds for these Fr\'echet classes. We show that the improved Fr\'echet--Hoeffding bounds are pointwise sharp for these classes in the presence of uncertainty in the marginals, while a counterexample yields that they are not pointwise sharp in the absence of uncertainty in the marginals, even in dimension 2. The latter result sheds new light on the improved Fr\'echet--Hoeffding bounds, since Tankov [30] has showed that, under certain conditions, these bounds are sharp in dimension 2.
We consider a nonlinear random walk which, in each time step, is free to choose its own transition probability within a neighborhood (w.r.t. Wasserstein distance) of the transition probability of a fixed L\'evy process. In analogy to the classical framework we show that, when passing from discrete to continuous time via a scaling limit, this nonlinear random walk gives rise to a nonlinear semigroup. We explicitly compute the generator of this semigroup and corresponding PDE as a perturbation of the generator of the initial L\'evy process.
We consider robust pricing and hedging for options written on multiple assets given market option prices for the individual assets. The resulting problem is called the multi-marginal martingale optimal transport problem. We propose two numerical methods to solve such problems: using discretisation and linear programming applied to the primal side and using penalisation and deep neural networks optimisation applied to the dual side. We prove convergence for our methods and compare their numerical performance. We show how adding further information about call option prices at additional maturities can be incorporated and narrows down the no-arbitrage pricing bounds. Finally, we obtain structural results for the case of the payoff given by a weighted sum of covariances between the assets.
We study a variant of the martingale optimal transport problem in a multi-period setting to derive robust price bounds on a financial derivative. On top of marginal and martingale constraints, we introduce a time-homogeneity assumption, which restricts the variability of the forward-looking transitions of the martingale across time. We provide a dual formulation in terms of superhedging and discuss relaxations of the time-homogeneity assumption by adding market frictions. In financial terms, the introduced time-homogeneity corresponds to a time-consistency condition for call prices, given the state of the stock. The time homogeneity assumption leads to improved price bounds since market data from many time points can be incorporated effectively. The approach is illustrated with two numerical examples.
We study the stability of entropically regularized optimal transport with respect to the marginals. Lipschitz continuity of the value and Hölder continuity of the optimal coupling in p-Wasserstein distance are obtained under general conditions including quadratic costs and unbounded marginals. The results for the value extend to regularization by an arbitrary divergence. Two techniques are presented: The first compares an optimal coupling with its so-called shadow, a coupling induced on other marginals by an explicit construction. The second transforms one set of marginals by a change of coordinates and thus reduces the comparison of differing marginals to the comparison of differing cost functions under the same marginals.
We study MinMax solution methods for a general class of optimization problems related to (and including) optimal transport. Theoretically, the focus is on fitting a large class of problems into a single MinMax framework and generalizing regularization techniques known from classical optimal transport. We show that regularization techniques justify the utilization of neural networks to solve such problems by proving approximation theorems and illustrating fundamental issues if no regularization is used. We further study the relation to the literature on generative adversarial nets, and analyze which algorithmic techniques used therein are particularly suitable to the class of problems studied in this paper. Several numerical experiments showcase the generality of the setting and highlight which theoretical insights are most beneficial in practice.
We consider settings in which the distribution of a multivariate random variable is partly ambiguous. We assume the ambiguity lies on the level of the dependence structure, and that the marginal distributions are known. Furthermore, a current best guess for the distribution, called reference measure, is available. We work with the set of distributions that are both close to the given reference measure in a transportation distance (e.g. the Wasserstein distance), and additionally have the correct marginal structure. The goal is to find upper and lower bounds for integrals of interest with respect to distributions in this set. The described problem appears naturally in the context of risk aggregation. When aggregating different risks, the marginal distributions of these risks are known and the task is to quantify their joint effect on a given system. This is typically done by applying a meaningful risk measure to the sum of the individual risks. For this purpose, the stochastic interdependencies between the risks need to be specified. In practice the models of this dependence structure are however subject to relatively high model ambiguity. The contribution of this paper is twofold: Firstly, we derive a dual representation of the considered problem and prove that strong duality holds. Secondly, we propose a generally applicable and computationally feasible method, which relies on neural networks, in order to numerically solve the derived dual problem. The latter method is tested on a number of toy examples, before it is finally applied to perform robust risk aggregation in a real world instance.
This note shows that, for a fixed Lipschitz constant $L > 0$, one layer neural networks that are $L$-Lipschitz are dense in the set of all $L$-Lipschitz functions with respect to the uniform norm on bounded sets.