We give a damped inexact Newton method for entropy-regularized least-squares on the nonnegative orthant that converges globally at a linear rate with O(^-1) iteration complexity, locally at a superlinear-to-quadratic rate, and is immune to the finite-precision overflow that limits classical dual solvers. A scale-shape decomposition of the primal – separating its scale from its direction – produces a dual with a nonsingular Jacobian. Objectives and Jacobians are evaluated through stable log-sum-exp and softmax primitives. Lambert W bounds on the scale uniformly control the Jacobian's spectrum, from which both rates follow. The solution map is jointly Lipschitz in the data, regularization parameter, and reference measure, and extends continuously to the vanishing-regularization limit. Experiments on a problem from analytic continuation of quantum Monte Carlo data confirm the predicted overflow resilience and convergence behavior.
In this paper, we study Lipschitz continuity of the solution mappings of regularized least-squares problems for which the convex regularizers have (Fenchel) conjugates that are C2-cone reducible. Our approach, by using Robinson's strong regularity on the dual problem, allows us to obtain new characterizations of Lipschitz stability that rely solely on firstorder information, thus bypassing the need to explore second-order information (curvature) of the regularizer. We show that these solution mappings are automatically Lipschitz continuous around the points in question whenever they are locally single-valued. We leverage our findings to obtain new characterizations of full stability and tilt stability for a broader class of convex additive-composite problems.
The Maximum Entropy on the Mean (MEM) method provides a flexible computational framework for solving inverse problems by combining data fidelity with entropy-based regularization. In practice, however, the prior distribution is typically unknown but can be estimated from data, giving rise to the empirical MEM method. We establish a parametric convergence rate of O(n^-1/2) in expectation for empirical MEM, improving upon the previously established O(n^-1/4) guarantee by King-Roskamp et al. (2026). Our proof is based on a novel stability analysis of the primal and dual optimization problems under perturbations of the underlying probability measure, relying only on foundational tools from convex analysis and probability. We further show that the MEM dual problem admits a reformulation as an expected risk minimization problem, thereby placing MEM within the modern framework of stochastic optimization and enabling scalable stochastic gradient algorithms for large-scale inverse problems. Together, these results place empirical MEM as a statistically and computationally efficient methodology for data-driven inverse problems.
We study perturbation and stability properties of solution mappings associated with convex regularized least-squares problems. We first establish an implicit function theorem for generalized equations governed by maximally monotone operators and smooth perturbations, yielding conditions for local Lipschitz continuity, directional differentiability, and semismoothness of solution mappings. We then characterize the kernel of generalized Hessians of convex functions through the subspace parallel to the subdifferential under C^2-cone reducibility assumptions on the conjugate of the regularizer, thereby replacing difficult second-order objects by tractable first-order conditions. As applications, we derive stability results for broad classes of regularized least-squares problems, including weighted polyhedral support-function regularizers and piecewise linear-quadratic penalties. Our framework unifies and extends several recent results on Lipschitz stability for LASSO-type models and monotone generalized equations.
We establish the theoretical framework for implementing the maximumn entropy on the mean (MEM) method for linear inverse problems in the setting of approximate (data-driven) priors. We prove a.s. convergence for empirical means and further develop general estimates for the difference between the MEM solutions with different priors μ and ν based upon the epigraphical distance between their respective log-moment generating functions. These estimates allow us to establish a rate of convergence in expectation for empirical means. We illustrate our results with denoising on MNIST and Fashion-MNIST data sets.
We explore a method of statistical estimation called Maximum Entropy on the Mean (MEM) which is based on an information-driven criterion that quantifies the compliance of a given point with a reference prior probability measure. At the core of this approach lies the MEM function which is a partial minimization of the Kullback-Leibler divergence over a linear constraint. In many cases, it is known that this function admits a simpler representation (known as the Cram\'er rate function). Via the connection to exponential families of probability distributions, we study general conditions under which this representation holds. We then address how the associated MEM estimator gives rise to a wide class of MEM-based regularized linear models for solving inverse problems. Finally, we propose an algorithmic framework to solve these problems efficiently based on the Bregman proximal gradient method, alongside proximal operators for commonly used reference distributions. The article is complemented by a software package for experimentation and exploration of the MEM approach in applications.
Submodular functions, defined on continuous or discrete domains, arise in numerous applications. We study the minimization of the difference of two submodular (DS) functions, over both domains, extending prior work restricted to set functions. We show that all functions on discrete domains and all smooth functions on continuous domains are DS. For discrete domains, we observe that DS minimization is equivalent to minimizing the difference of two convex (DC) functions, as in the set function case. We propose a novel variant of the DC Algorithm (DCA) and apply it to the resulting DC Program, obtaining comparable theoretical guarantees as in the set function case. The algorithm can be applied to continuous domains via discretization. Experiments demonstrate that our method outperforms baselines in integer compressive sensing and integer least squares.
This paper is concerned with augmented Lagrangian methods for the treatment of fully convex composite optimization problems. We extend the classical relationship between augmented Lagrangian methods and the proximal point algorithm to the inexact and safeguarded scheme in order to state global primal-dual convergence results. Our analysis distinguishes the regular case, where a stationary minimizer exists, and the irregular case, where all minimizers are nonstationary. Furthermore, we suggest an elastic modification of the standard safeguarding scheme which preserves primal convergence properties while guaranteeing convergence of the dual sequence to a multiplier in the regular situation. Although important for nonconvex problems, the standard safeguarding mechanism leads to weaker convergence guarantees for convex problems than the classical augmented Lagrangian method. Our elastic safeguarding scheme combines the advantages of both while avoiding their shortcomings.
Many fields of physics use quantum Monte Carlo techniques, but struggle to estimate dynamic spectra via the analytic continuation of imaginary-time quantum Monte Carlo data. One of the most ubiquitous approaches to analytic continuation is the maximum entropy method (MEM). We supply a dual Newton optimization algorithm to be used within the MEM and provide analytic bounds for the algorithm's error. The optimization algorithm is freely availible on github (repository). The MEM is typically used with Bryan's controversial algorithm (Rothkopf 2020 Data 5 55). We present new theoretical issues that are not yet in the literature. Our algorithm has all the theoretical benefits of Bryan's algorithm without these theoretical issues the implementation of the dual Newton optimizer within the MEM is freely available on github (repository). We compare the MEM with Bryan's optimization to the MEM with our dual Newton optimization on test problems from lattice quantum chromodynamics and plasma physics. These comparisons show that in the presence of noise the dual Newton algorithm produces better estimates and error bars; this indicates the limits of Bryan's algorithm's applicability. We use the MEM to investigate authentic quantum Monte Carlo data for the uniform electron gas at warm dense matter conditions and further substantiate the roton-type feature in the dispersion relation.
Understanding the dynamic properties of uniform electron gas (UEG) is important for numerous applicationsranging from semiconductor physics to exotic warm dense matter. In this work, we apply the maximum entropymethod (MEM), as implemented by Thomas Chunaet al.,J. Phys. A Math. Theor.58, 335203 (2025),toab initiopath integral Monte Carlo (PIMC) results for the imaginary-time correlation functionF(q,tau) to estimate thedynamic structure factorS(q,omega) over an unprecedented range of densities at the electronic Fermi temperature.To conduct the MEM, we propose to construct the Bayesian prior mu from the PIMC data. Constructing thestatic approximation leads to a drastic improvement inS(q,omega) estimate over using the simpler random phaseapproximation (RPA) as the Bayesian prior. We present results for the strongly coupled electron liquid regimewithrs=50,...,200, which reveal a pronounced roton-type feature and an incipient double peak structure inS(q,omega) for intermediate wave numbers atrs=200. We also find that our dynamic structure factors satisfy knownsum rules, even though these sum rules are not enforced explicitly. To verify our results, we show that our MEM estimates converge to the RPA limit at higher densities and weaker couplingrs=2 and 5. Further, we compare with two different existing results at intermediate density and coupling strengthrs=10 and 20, and we find good agreement with more conservative estimates. Combining all of our results forrs=2,5,10,20,50,100,200, we present estimates of a dispersion relation that show a continuous deepening of its minimum value at higher coupling. An advantage of our setup is that it is not specific to the UEG, thereby opening up new avenues to study the dynamics of real warm dense matter systems based on cutting-edge PIMC simulations in future works.
This paper studies well-posedness and parameter sensitivity of the Square Root LASSO (SR-LASSO), an optimization model for recovering sparse solutions to linear inverse problems in finite dimension. An advantage of the SR-LASSO (e.g., over the standard LASSO) is that the optimal tuning of the regularization parameter is robust with respect to measurement noise. This paper provides three point-based regularity conditions at a solution of the SR-LASSO: the weak, intermediate, and strong assumptions. It is shown that the weak assumption implies uniqueness of the solution in question. The intermediate assumption yields a directionally differentiable and locally Lipschitz solution map (with explicit Lipschitz bounds), whereas the strong assumption gives continuous differentiability of said map around the point in question. Our analysis leads to new theoretical insights on the comparison between SR-LASSO and LASSO from the viewpoint of tuning parameter sensitivity: noise-robust optimal parameter choice for SR-LASSO comes at the "price" of elevated tuning parameter sensitivity. Numerical results support and showcase the theoretical findings.
In this paper we prove the existence of Hölder continuous terminal embeddings of any desired X ⊆ℝ^d into ℝ^m with m=𝒪(ε^-2ω(S_X)^2), for arbitrarily small distortion ε, where ω(S_X) denotes the Gaussian width of the unit secants of X. More specifically, when X is a finite set we provide terminal embeddings that are locally 1/2-Hölder almost everywhere, and when X is infinite with positive reach we give terminal embeddings that are locally 1/4-Hölder everywhere sufficiently close to X (i.e., within all tubes around X of radius less than X's reach). When X is a compact d-dimensional submanifold of ℝ^N, an application of our main results provides terminal embeddings into 𝒪̃(d)-dimensional space that are locally Hölder everywhere sufficiently close to the manifold.
Minimizing the difference of two submodular (DS) functions is a problem that naturally occurs in various machine learning problems. Although it is well known that a DS problem can be equivalently formulated as the minimization of the difference of two convex (DC) functions, existing algorithms do not fully exploit this connection. A classical algorithm for DC problems is called the DC algorithm (DCA). We introduce variants of DCA and its complete form (CDCA) that we apply to the DC program corresponding to DS minimization. We extend existing convergence properties of DCA, and connect them to convergence properties on the DS problem. Our results on DCA match the theoretical guarantees satisfied by existing DS algorithms, while providing a more complete characterization of convergence properties. In the case of CDCA, we obtain a stronger local minimality guarantee. Our numerical results show that our proposed algorithms outperform existing baselines on two applications: speech corpus selection and feature selection.
The projection onto the epigraph or a level set of a closed proper convex function can be achieved by finding a root of a scalar equation that involves the proximal operator as a function of the proximal parameter. This paper develops the variational analysis of this scalar equation. The approach is based on a study of the variational-analytic properties of general convex optimization problems that are (partial) infimal projections of the sum of the function in question and the perspective map of a convex kernel. When the kernel is the Euclidean norm squared, the solution map corresponds to the proximal map, and thus, the variational properties derived for the general case apply to the proximal case. Properties of the value function and the corresponding solution map—including local Lipschitz continuity, directional differentiability, and semismoothness—are derived. An SC 1 optimization framework for computing epigraphical and level-set projections is, thus, established. Numerical experiments on one-norm projection illustrate the effectiveness of the approach as compared with specialized algorithms. Funding: This work was supported by Natural Sciences and Engineering Research Council of Canada (NSERC). M.P. Friedlander was supported by NSERC discovery grant [Grant RGPIN-2017-04461]. T. Hoheisel was supported by NSERC discovery grant [Grant RGPIN-2017-04035]. A. Goodwin’s work was partially supported by an NSERC summer research stipend.
This paper provides a variational analysis of the unconstrained formulation of the LASSO problem, ubiquitous in statistical learning, signal processing, and inverse problems. In particular, we establish smoothness results for the optimal value as well as Lipschitz properties of the optimal solution as functions of the right-hand side (or measurement vector) and the regularization parameter. Moreover, we show how to apply the proposed variational analysis to study the sensitivity of the optimal solution to the tuning parameter in the context of compressed sensing with subgaussian measurements. Our theoretical findings are validated by numerical experiments.
Mathematical programs with vanishing constraints (MPVCs) are a class of nonlinear optimization problems with applications to various engineering problems such as truss topology design and robot motion planning. MPVCs are difficult problems from both a theoretical and numerical perspective: the combinatorial nature of the vanishing constraints often prevents standard constraint qualifications and optimality conditions from being attained; moreover, the feasible set is inherently nonconvex, and often has no interior around points of interest. In this paper, we therefore study and compare four regularization methods for the numerical solution of MPVCS. Each method depends on a single regularization parameter, which is used to embed the original MPVC into a sequence of standard nonlinear programs. Convergence results for these methods based on both exact and approximate stationary of the subproblems are established under weak assumptions. The improved regularity of the subproblems is studied by providing sufficient conditions for the existence of KKT multipliers. Numerical experiments, based on applications in truss topology design and an optimal control problem from aerothermodynamics, complement the theoretical analysis and comparison of the regularization methods. The computational results highlight the benefit of using regularization over applying a standard solver directly, and they allow us to identify two promising regularization schemes.
In this paper, we provide a full conjugacy and subdifferential calculus for convex convex-composite functions in finite-dimensional space. Our approach, based on infimal convolution and cone convexity, is straightforward. The results are established under a verifiable Slater-type condition, with relaxed monotonicity and without lower semicontinuity assumptions on the functions in play. The versatility of our findings is illustrated by a series of applications in optimization and matrix analysis, including conic programming, matrix-fractional, variational Gram, and spectral functions.
Deep neural networks perform well on real world data but are prone to adversarial perturbations: small changes in the input easily lead to misclassification. In this work, we propose an attack methodology not only for cases where the perturbations are measured by $\ell_p$ norms, but in fact any adversarial dissimilarity metric with a closed proximal form. This includes, but is not limited to, $\ell_1, \ell_2$, and $\ell_\infty$ perturbations; the $\ell_0$ counting norm (i.e. true sparseness); and the total variation seminorm, which is a (non-$\ell_p$) convolutional dissimilarity measuring local pixel changes. Our approach is a natural extension of a recent adversarial attack method, and eliminates the differentiability requirement of the metric. We demonstrate our algorithm, ProxLogBarrier, on the MNIST, CIFAR10, and ImageNet-1k datasets. We consider undefended and defended models, and show that our algorithm easily transfers to various datasets. We observe that ProxLogBarrier outperforms a host of modern adversarial attacks specialized for the $\ell_0$ case. Moreover, by altering images in the total variation seminorm, we shed light on a new class of perturbations that exploit neighboring pixel information.
Image deblurring is a notoriously challenging ill-posed inverse problem. In recent years, a wide variety of approaches have been proposed based upon regularization at the level of the image or on techniques from machine learning. We propose an alternative approach, shifting the paradigm towards regularization at the level of the probability distribution on the space of images. Our method is based upon the idea of maximum entropy on the mean wherein we work at the level of the probability density function of the image whose expectation is our estimate of the ground truth. Using techniques from convex analysis and probability theory, we show that the method is computationally feasible and amenable to very large blurs. Moreover, when images are imbedded with symbology (a known pattern), we show how our method can be applied to approximate the unknown blur kernel with remarkable effects. While our method is stable with respect to small amounts of noise, it does not actively denoise. However, for moderate to large amounts of noise, it performs well by preconditioned denoising with a state of the art method.
Florian Jarre合作论文数Mathematisches Institut
Lehrstuhl für Mathematische Optimierung
Universitätsstraße 12