The Sliced Wasserstein (SW) distance has emerged as a computationally attractive alternative to the Wasserstein distance by leveraging one-dimensional optimal transport along random projections. Standard estimators of the SW distance rely on Monte Carlo averages of one-dimensional Wasserstein distances computed via quantile functions, which require sorting projected samples and access to full datasets. In this work, we introduce a new class of estimators for the Sliced Wasserstein distance based on cumulative distribution functions (CDFs) of projected measures, that avoid sorting and scale via massive dataset parallelism. This class includes several estimators, some of them being indexed by hyperparameters controlling their variance or smoothness. We show that they are especially well suited to scenarios in which CDFs are more tractable than quantile functions, such as mixtures of Gaussians, and moreover that they are also naturally compatible with federated learning, since CDFs of projected data can be computed and aggregated locally without requiring the exchange of raw samples.
The stability of solutions to optimal transport problems under variation of the measures is fundamental from a mathematical viewpoint: it is closely related to the convergence of numerical approaches to solve optimal transport problems and justifies many of the applications of optimal transport. In this article, we introduce the notion of strong c-concavity, and we show that it plays an important role for proving stability results in optimal transport for general cost functions c. We then introduce a differential criterion for proving that a function is strongly c-concave, under an hypothesis on the cost introduced originally by Ma-Trudinger-Wang for establishing regularity of optimal transport maps. Finally, we provide two examples where this stability result can be applied, for cost functions taking value +$\infty$ on the sphere: the reflector problem and the Gaussian curvature measure prescription problem.
We study a nonlinear multimarginal optimal transport problem arising in risk management, where the objective is to maximize a spectral risk measure of the pushforward of a coupling by a cost function. Although this problem is inherently nonlinear, it is known to have an equivalent linear reformulation as a multimarginal transport problem with an additional marginal. We introduce a Lagrangian particle discretization of this problem, in which admissible couplings are approximated by uniformly weighted point clouds, and marginal constraints are enforced through Wasserstein penalization. We prove quantitative convergence results for this discretization as the number of particles tends to infinity. The convergence rate is shown to be governed by the uniform quantization error of an optimal solution, and can be bounded in terms of the geometric properties of its support, notably its box dimension. In the case of univariate marginals and supermodular cost functions, where optimal couplings are known to be comonotone, we obtain sharper convergence rates expressed in terms of the asymptotic quantization errors of the marginals themselves. We also discuss the particular case of conditional value at risk, for which the problem reduces to a multimarginal partial transport formulation. Finally, we illustrate our approach with numerical experiments in several application domains, including risk management and partial barycenters, as well as some artificial examples with a repulsive cost.
We prove quantitative bounds on the stability of optimal transport maps and Kantorovich potentials from a fixed source measure ρ under variations of the target measure μ, when the cost function is the squared Riemannian distance on a Riemannian manifold. Previous works were restricted to subsets of Euclidean spaces, or made specific assumptions either on the manifold, or on the regularity of the transport maps. Our proof techniques combine entropy-regularized optimal transport with spectral and integral-geometric techniques. As some of the arguments do not rely on the Riemannian structure, our work also paves the way towards understanding stability of optimal transport in more general geometric spaces.
Sliced Wasserstein distances are widely used in practice as a computationally efficient alternative to Wasserstein distances in high dimensions. In this paper, motivated by theoretical foundations of this alternative, we prove quantitative estimates between the sliced 1-Wasserstein distance and the 1-Wasserstein distance. We construct a concrete example to demonstrate the exponents in the estimate is sharp. We also provide a general analysis for the case where slicing involves projections onto k-planes and not just lines.
We establish quantitative stability bounds for the quadratic optimal transport map T_μ between a fixed probability density ρ and a probability measure μ on ℝ^d. Under general assumptions on ρ, we prove that the map μ↦ T_μ is bi-Hölder continuous, with dimension-free Hölder exponents. The linearized optimal transport metric W_2,ρ(μ,ν)=T_μ-T_ν_L^2(ρ) is therefore bi-Hölder equivalent to the 2-Wasserstein distance, which justifies its use in applications. We show this property in the following cases: (i) for any log-concave density ρ with full support in ℝ^d, and any log-bounded perturbation thereof; (ii) for ρ bounded away from 0 and +∞ on a John domain (e.g., on a bounded Lipschitz domain), while the only previously known result of this type assumed convexity of the domain; (iii) for some important families of probability densities on bounded domains which decay or blow-up polynomially near the boundary. Concerning the sharpness of point (ii), we also provide examples of non-John domains for which the Brenier potentials do not satisfy any Hölder stability estimate. Our proofs rely on local variance inequalities for the Brenier potentials in small convex subsets of the support of ρ, which are glued together to deduce a global variance inequality. This gluing argument is based on two different strategies of independent interest: one of them leverages the properties of the Whitney decomposition in bounded domains, the other one relies on spectral graph theory.
In this article, we propose a numerical method to solve semi-discrete optimal transport problems for gigantic pointsets (108 points and more). By pushing the limits by several orders of magnitude, it opens the path to new applications in cosmology, fluid simulation and data science to name but a few. The method is based on a new algorithm that computes (generalized) Voronoi diagrams in parallel and in a distributed way. First we make the simple observation that the cells defined by a subgraph of the Delaunay graph contain the Voronoi cells, and that one can deduce the missing edges from the intersections between those cells. Based on this observation, we introduce the Distributed Voronoi Diagram algorithm (DVD) that can be used on a cluster and that exchanges vertices between the nodes as need be. We also report early experimental results, demonstrating that the DVD algorithm has the potential to solve some giga-scale semi-discrete optimal transport problems encountered in computational cosmology.
In this paper, we investigate the properties of the Sliced Wasserstein Distance (SW) when employed as an objective functional. The SW metric has gained significant interest in the optimal transport and machine learning literature, due to its ability to capture intricate geometric properties of probability distributions while remaining computationally tractable, making it a valuable tool for various applications, including generative modeling and domain adaptation. Our study aims to provide a rigorous analysis of the critical points arising from the optimization of the SW objective. By computing explicit perturbations, we establish that stable critical points of SW cannot concentrate on segments. This stability analysis is crucial for understanding the behaviour of optimization algorithms for models trained using the SW objective. Furthermore, we investigate the properties of the SW objective, shedding light on the existence and convergence behavior of critical points. We illustrate our theoretical results through numerical experiments.
We study the problem of maximizing a spectral risk measure of a given output function which depends on several underlying variables, whose individual distributions are known but whose joint distribution is not. We establish and exploit an equivalence between this problem and a multi-marginal optimal transport problem. We use this reformulation to establish explicit, closed form solutions when the underlying variables are one dimensional, for a large class of output functions. For higher dimensional underlying variables, we identify conditions on the output function and marginal distributions under which solutions concentrate on graphs over the first variable and are unique, and, for general output functions, we find upper bounds on the dimension of the support of the solution. We also establish a stability result on the maximal value and maximizing joint distributions when the output function, marginal distributions and spectral function are perturbed; in addition, when the variables one dimensional, we show that the optimal value exhibits Lipschitz dependence on the marginal distributions for a certain class of output functions. Finally, we show that the equivalence to a multi-marginal optimal transport problem extends to maximal correlation measures of multi-dimensional risks; in this setting, we again establish conditions under which the solution concentrates on a graph over the first marginal.
We study the quantitative stability of the mapping that to a measure associates its pushforward measure by a fixed (non-smooth) optimal transport map. We exhibit a tight Hölder-behavior for this operation under minimal assumptions. Our proof essentially relies on a new bound that quantifies the size of the singular sets of a convex and Lipschitz continuous function on a bounded domain.
Wasserstein barycenters define averages of probability measures in a geometrically meaningful way. Their use is increasingly popular in applied fields, such as image, geometry or language processing. In these fields however, the probability measures of interest are often not accessible in their entirety and the practitioner may have to deal with statistical or computational approximations instead. In this article, we quantify the effect of such approximations on the corresponding barycenters. We show that Wasserstein barycenters depend in a H{\"o}lder-continuous way on their marginals under relatively mild assumptions. Our proof relies on recent estimates that quantify the strong convexity of the dual quadratic optimal transport problem and a new result that allows to control the modulus of continuity of the push-forward operation under a (not necessarily smooth) optimal transport map.
This work studies the quantitative stability of the quadratic optimal transport map between a fixed probability density $\rho$ and a probability measure $\mu$ on R^d , which we denote T$\mu$. Assuming that the source density $\rho$ is bounded from above and below on a compact convex set, we prove that the map $\mu$ $\rightarrow$ T$\mu$ is bi-H{\"o}lder continuous on large families of probability measures, such as the set of probability measures whose moment of order p > d is bounded by some constant. These stability estimates show that the linearized optimal transport metric W2,$\rho$($\mu$, $\nu$) = T$\mu$ -- T$\nu$ L 2 ($\rho$,R d) is bi-H{\"o}lder equivalent to the 2-Wasserstein distance on such sets, justifiying its use in applications.
When expressed in Lagrangian variables, the equations of motion for compressible (barotropic) fluids have the structure of a classical Hamiltonian system in which the potential energy is given by the internal energy of the fluid. The dissipative counterpart of such a system coincides with the porous medium equation, which can be cast in the form of a gradient flow for the same internal energy. Motivated by these related variational structures, we propose a particle method for both problems in which the internal energy is replaced by its Moreau-Yosida regularization in the L2 sense, which can be efficiently computed as a semi-discrete optimal transport problem. Using a modulated energy argument which exploits the convexity of the problem in Eulerian variables, we prove quantitative convergence estimates towards smooth solutions. We verify such estimates by means of several numerical tests.
Is it possible to shape a piece of glass so that it refracts and concentrates sunlight in order to produce a given image? The modelling of this kind of problem leads to nonlinear second-order partial differential equations, which belong to the family of Monge–Ampère equations. We will see how semi-discrete methods, that can be traced back to Minkowski’s works, allow us to numerically solve such equations.
Several issues in machine learning and inverse problems require to generate discrete data, as if sampled from a model probability distribution. A common way to do so relies on the construction of a uniform probability distribution over a set of $N$ points which minimizes the Wasserstein distance to the model distribution. This minimization problem, where the unknowns are the positions of the atoms, is non-convex. Yet, in most cases, a suitably adjusted version of Lloyd's algorithm -- in which Voronoi cells are replaced by Power cells -- leads to configurations with small Wasserstein error. This is surprising because, again, of the non-convex nature of the problem, as well as the existence of spurious critical points. We provide explicit upper bounds for the convergence speed of this Lloyd-type algorithm, starting from a cloud of points sufficiently far from each other. This already works after one step of the iteration procedure, and similar bounds can be deduced, for the corresponding gradient descent. These bounds naturally lead to a modified Poliak-Lojasiewicz inequality for the Wasserstein distance cost, with an error term depending on the distances between Dirac masses in the discrete distribution.
This paper provides a theoretical and numerical approach to show existence, uniqueness, and the numerical determination of metalenses refracting radiation with energy patterns. The theoretical part uses ideas from optimal transport and for the numerical solution we study and implement a damped Newton algorithm to solve the semi discrete problem. A detailed analysis is carried out to solve the near field one source refraction problem and extensions to the far field are also mentioned.
This chapter describes techniques for the numerical resolution of optimal transport problems. We will consider several discretizations of these problems, and we will put a strong focus on the mathematical analysis of the algorithms to solve the discretized problems. We will describe in detail the following discretizations and corresponding algorithms: the assignment problem and Bertsekas auction's algorithm; the entropic regularization and Sinkhorn-Knopp's algorithm; semi-discrete optimal transport and Oliker-Prussner or damped Newton's algorithm, and finally semi-discrete entropic regularization. Our presentation highlights the similarity between these algorithms and their connection with the theory of Kantorovich duality.
We study a model of crowd motion following a gradient vector field, with possibly additional interaction terms such as attraction/repulsion, and we present a numerical scheme for its solution through a Lagrangian discretization. The density constraint of the resulting particles is enforced by means of a partial optimal transport problem at each time step. We prove the convergence of the discrete measures to a solution of the continuous PDE describing the crowd motion in dimension one. In a second part, we show how a similar approach can be used to construct a Lagrangian discretization of a linear advection-diffusion equation, interpreted as a gradient flow in Wasserstein space. We provide also a numerical implementation in 2D to demonstrate the feasibility of the computations.
This work studies an explicit embedding of the set of probability measures into a Hilbert space, defined using optimal transport maps from a reference probability density. This embedding linearizes to some extent the 2-Wasserstein space, and enables the direct use of generic supervised and unsupervised learning algorithms on measure data. Our main result is that the embedding is (bi-)H\older continuous, when the reference density is uniform over a convex set, and can be equivalently phrased as a dimension-independent H\older-stability results for optimal transport maps.
M. Yvinec合作论文数Unit?? de Sophia Antipolis,
Projet GEOMETRICA2