In this paper, we propose an Anderson-accelerated stochastic extragradient algorithm for solving a class of stochastic variational inequalities, by incorporating Anderson acceleration into the stochastic extragradient method under a stochastic approximation framework. A key challenge in our setting is that the pseudomonotonicity assumption is only imposed on the expectation of the stochastic operator, rather than on the individual stochastic operator itself and the sample averages utilized in the algorithm. We prove that, despite the lack of pseudomonotonicity in the sampled operators, the sequence generated by the proposed algorithm converges almost surely to a solution of the stochastic variational inequality problem. Additionally, we establish the sublinear convergence rate of the proposed algorithm in terms of the mean residual function, along with its optimal oracle complexity. Finally, we validate the effectiveness of the proposed algorithm through numerical experiments.
In this paper, we study a class of constrained group sparse ℓ _0 regularized optimization problems, where the loss function is convex but nonsmooth and the feasible set is defined by box constraints. First, we propose a smoothing proximal gradient block-coordinate (SPGBC) algorithm, which is a novel combination of the proximal gradient block-coordinate algorithm and the smoothing method. We prove that any accumulation point of the iterates generated by it is a local minimizer of the considered problem and its zero entries can be identified in finite iterations. Moreover, we show that the proposed SPGBC algorithm achieves a local convergence rate of 𝒪(k^-(1-ν )) on the objective function value, where ν∈ (1/2,1) comes from the decay exponent of the smoothing parameter. Second, we consider a randomized variant of the SPGBC algorithm, the R-SPGBC algorithm, and obtain that the iterates generated by it converge to a subset of local minimizers of the original problem with probability 1. In addition, we establish that the R-SPGBC algorithm attains a sublinear convergence rate in expectation. Finally, some numerical examples are performed to show the efficiency of the proposed algorithms.
Robust principal component analysis is an important representative method in data analysis. It is usually viewed as an optimization problem involving the rank and ℓ_0-norm of matrices. In this paper, we study the rank and ℓ_0 regularized optimization problem and its matrix factorization problem. We establish their equivalences on global minimizers and stationary points, respectively. Furthermore, we construct a broadly applicable equivalent nonconvex relaxation framework for the constrained factorization model in the sense of global minimizers and stationary points with strong optimality conditions (called strong stationary points). For the general factorization problem with lower semicontinuous regularizers and a loss function whose gradient is locally Lipschitz, we propose a novel proximal gradient-based algorithm based on joint and alternating calculation with convergence to its limiting-critical points. The algorithm can attain the stationary points of the original problem and its adaptive counterpart can attain the strong stationary points of the factorization problem.
Balancing spectral, spatial, and temporal resolutions is a key challenge in spectral imaging. The Dual-Camera Coded Aperture Snapshot Spectral Imaging (DC-CASSI) system alleviates this trade-off but suffers from severely ill-posed reconstruction problems due to its high compression ratio. Existing methods are constrained by scene-specific tuning or excessive reliance on paired training data. To address these issues, we propose a Total Variation (TV) subgradient-guided multi-source fusion framework for DC-CASSI reconstruction, comprising three core components: (1) An end-to-end Single-Disperser CASSI (SD-CASSI) observation model based on the tensor-form Kronecker δ, which establishes a rigorous mathematical foundation for physical constraints while enabling efficient adjoint operator implementation; (2) An adaptive spatial reference generator that integrates SD-CASSI's physical model and RGB subspace constraint, generating the reference image as reliable spatial prior; (3) A TV subgradient-guided regularization term that encodes local structural directions from the reference image into spectral reconstruction, achieving high-quality fused results. The framework is validated on simulated datasets and real-world datasets. Experimental results demonstrate that it achieves state-of-the-art reconstruction performance and robust noise resilience. This work not only establishes an interpretable theoretical foundation for subgradient-guided fusion but also provides a practical fusion-based paradigm for high-fidelity spectral image reconstruction in DC-CASSI systems. Source code: https://github.com/bestwishes43/ADMM-TVDS.
In this paper, we study a class of structured sparse optimization problems characterized by a convex, possibly non-smooth loss function and a capped- ℓ _1 penalty. This model can provide an exact continuous relaxation for the problems with a cardinality penalty. First, we propose an Accelerated Nested Proximal Gradient (ANPG) algorithm, which employs a nested proximal structure and selective extrapolation to enhance computational efficiency. Under some mild and easily verifiable conditions, we prove the subsequence convergence of the ANPG algorithm to the lifted stationary points of the considered problem, which correspond to the strong local minimizers of the corresponding cardinality penalty problems. Moreover, we establish an O(k^-1) convergence rate on the objective function values, while a refined extrapolation strategy ensures sequence convergence of the iterates, albeit with a potential reduction in the theoretical convergence rate on the objective values. Second, for the cases with a smooth loss function, we further propose an Accelerated Proximal Gradient (APG) algorithm that guarantees the sequence convergence on the iterates and achieves a faster convergence rate of o(k^-2) in terms of objective values. Finally, the effectiveness of the ANPG and APG algorithms, as well as the high quality of the solutions they produce, are verified through numerical experiments on the least absolute deviation regression problem and the sparse logistic regression problem, respectively.
Lately, a novel swarm intelligence model, namely the consensus-based optimization (CBO) algorithm, was introduced to deal with the global optimization problems. Limited by the conditions of Ito's formula, the convergence analysis of the previous CBO finite particle system mainly focuses on the problem with smooth objective function. With the help of smoothing method, this paper achieves a breakthrough by proposing an effective CBO algorithm for solving the global solution of a nonconvex, nonsmooth, and possible non-Lipschitz continuous minimization problem with theoretical analysis, which dose not rely on the mean-field limit. We indicate that the proposed algorithm exhibits a global consensus and converges to a common state with any initial data. Then, we give a more detailed error estimation on the objective function values along the state of the proposed algorithm towards the global minimum. Finally, some numerical examples are presented to illustrate the appreciable performance of the proposed method on solving the nonsmooth, nonconvex minimization problems.
In this paper, we focus on finding the global minimizer of a general unconstrained nonsmooth nonconvex optimization problem. Taking advantage of the smoothing method and the consensus-based optimization (CBO) method, we propose a novel smoothing iterative consensus-based optimization (SICBO) algorithm. First, we prove that the solution process of the proposed algorithm here exponentially converges to a common stochastic consensus point almost surely. Second, we establish a detailed theoretical analysis to ensure the small enough error between the objective function value at the consensus point and the optimal function value, to the best of our knowledge, which provides the first theoretical guarantee to the global optimality of the proposed algorithm for nonconvex optimization problems. Moreover, unlike the previously introduced CBO methods, the theoretical results are valid for the cases that the objective function is nonsmooth, nonconvex and perhaps non-Lipschitz continuous. Finally, several numerical examples are performed to illustrate the effectiveness of our proposed algorithm for solving the global minimizer of the nonsmooth and nonconvex optimization problems.
This paper proposes an extra gradient Anderson-accelerated algorithm for solving pseudomonotone variational inequalities, which uses the extra gradient scheme with line search to guarantee the global convergence and Anderson acceleration to have fast convergent rate. We prove that the sequence generated by the proposed algorithm from any initial point converges to a solution of the pseudomonotone variational inequality problem without assuming the Lipschitz continuity and contractive condition, which are used for convergence analysis of the extra gradient method and Anderson-accelerated method, respectively in existing literatures. Numerical experiments, particular emphasis on Harker-Pang problems, fractional programming problems, nonlinear complementarity problems, partial differential equation problems with free boundary and linear complementarity problems, are conducted to validate the effectiveness and good performance of the proposed algorithm comparing with the extra gradient method and Anderson-accelerated method.
In this paper, we are interested in finding the global minimizer of a nonsmooth nonconvex unconstrained optimization problem. By combining the discrete consensus-based optimization (CBO) algorithm and the gradient descent method, we develop a novel CBO algorithm with an extra gradient descent scheme evaluated by the forward-difference technique on the function values, where only the objective function values are used in the proposed algorithm. First, we prove that the proposed algorithm can exhibit global consensus in an exponential rate in two senses and possess a unique global consensus point. Second, we evaluate the error estimate between the objective function value on the global consensus point and its global minimum. In particular, as the parameter β tends to ∞, the error converges to zero and the convergence rate is 𝒪(logβ/β). Third, under some suitable assumptions on the objective function, we provide the number of iterations required for the mean square error in expectation to reach the desired accuracy. It is worth underlining that the theoretical analysis in this paper does not use the mean-field limit. Finally, we illustrate the improved efficiency and promising performance of our novel CBO method through some experiments on several nonconvex benchmark problems and the application to train deep neural networks.
For a class of sparse optimization problems with the penalty function of ‖ (· )_+‖ _0 , we first characterize its local minimizers and then propose an extrapolated hard thresholding algorithm to solve such problems. We show that the iterates generated by the proposed algorithm with ϵ >0 (where ϵ is the dry friction coefficient) have finite length, without relying on the Kurdyka-Łojasiewicz inequality. Furthermore, we demonstrate that the algorithm converges to an ϵ -local minimizer of this problem. For the special case that ϵ =0 , we establish that any accumulation point of the iterates is a local minimizer of the problem. Additionally, we analyze the convergence when an error term is present in the algorithm, showing that the algorithm still converges in the same manner as before, provided that the errors asymptotically approach zero. Finally, we conduct numerical experiments to verify the theoretical results of the proposed algorithm.
In this paper, we consider the Anderson acceleration method for solving the contractive fixed point problem, which is nonsmooth in general. We define a class of smoothing functions for the original nonsmooth fixed point mapping, which can be easily formulated for many cases (see section3). Then, taking advantage of the Anderson acceleration method, we proposed a Smoothing Anderson(m) algorithm, in which we utilized a smoothing function of the original nonsmooth fixed point mapping and update the smoothing parameter adaptively. In theory, we first demonstrate the r-linear convergence of the proposed Smoothing Anderson(m) algorithm for solving the considered nonsmooth contractive fixed point problem with r-factor no larger than c, where c is the contractive factor of the fixed point mapping. Second, we establish that both of the Smoothing Anderson(1) and the Smoothing EDIIS(1) algorithms are q-linear convergent with q-factor no larger than c. Finally, we present three numerical examples with practical applications from elastic net regression, free boundary problems for infinite journal bearings and non-negative least squares problem to illustrate the better performance of the proposed Smoothing Anderson(m) algorithm comparing with some popular methods.
In this paper, we focus on a class of convexly constrained nonsmooth convex–concave saddle point problems with cardinality penalties. Although such nonsmooth nonconvex–nonconcave and discontinuous min–max problems may not have a saddle point, we show that they have a local saddle point and a global minimax point, and some local saddle points have the lower bound properties. We define a class of strong local saddle points based on the lower bound properties for stability of variable selection. Moreover, we give a framework to construct continuous relaxations of the discontinuous min–max problems based on convolution, such that they have the same saddle points with the original problem. We also establish the relations between the continuous relaxation problems and the original problems regarding local saddle points, global minimax points, local minimax points and stationary points. Finally, we illustrate our results with distributionally robust sparse convex regression, sparse robust bond portfolio construction and sparse convex–concave logistic regression saddle point problems.
Rank regularized minimization problem is an ideal model for the low-rank matrix completion/recovery problem. The matrix factorization approach can transform the high-dimensional rank regularized problem to a low-dimensional factorized column-sparse regularized problem. The latter can greatly facilitate fast computations in applicable algorithms, but needs to overcome the simultaneous non-convexity of the loss and regularization functions. In this paper, we consider the factorized column-sparse regularized model. Firstly, we optimize this model with bound constraints, and establish a certain equivalence between the optimized factorization problem and rank regularized problem. Further, we strengthen the optimality condition for stationary points of the factorization problem and define the notion of strong stationary point. Moreover, we establish the equivalence between the factorization problem and its a nonconvex relaxation in the sense of global minimizers and strong stationary points. To solve the factorization problem, we design two types of algorithms and give an adaptive method to reduce their computation. The first algorithm is from the relaxation point of view and its iterates own some properties from global minimizers of the factorization problem after finite iterations. We give some analysis on the convergence of its iterates to the strong stationary point. The second algorithm is designed for directly solving the factorization problem. We improve the PALM algorithm introduced by Bolte et al. (Math Program Ser A 146:459-494, 2014) for the factorization problem and give its improved convergence results. Finally, we conduct numerical experiments to show the promising performance of the proposed model and algorithms for low-rank matrix completion.
In this paper, we propose a smoothing randomized block-coordinate proximal gradient (S-RBCPG) algorithm and a Bregman randomized block-coordinate proximal gradient (B-RBCPG) algorithm for minimizing the sum of two nonconvex nonsmooth functions, one of which is block separable. The pivotal tool of our analysis is the connection of the proximal gradient mapping with V-proximal mapping and Bregman proximal mapping. The S-RBCPG algorithm overcomes the non-smoothness of the objective function by utilizing the smoothing technique and we establish its subsequential convergence. Further, the B-RBCPG algorithm is designed for the case where the separable function is relatively smooth (that is, each separation part is relatively smooth). Then, we establish the ℛ -linear convergence rate of the B-RBCPG algorithm under expectation by assuming the Kurdyka-Łojasiewicz property on the objective function. Finally, we use some numerical experiments to illustrate the effectiveness and convergence of the proposed algorithms.
Matching is an important prerequisite for point clouds registration, which is to establish a reliable correspondence between two point clouds. This paper aims to improve recent theoretical and algorithmic results on discrete optimal transport (DOT), since it lacks robustness for the point clouds matching problems with large-scale affine or even nonlinear transformation. We first consider the importance of the used prior probability for accurate matching and give some theoretical analysis. Then, to solve the point clouds matching problems with complex deformation and noise, we propose an improved DOT model, which introduces an orthogonal matrix and a diagonal matrix into the classical DOT model. To enhance its capability of dealing with cases with outliers, we further bring forward a relaxed and regularized DOT model. Meantime, we propose two algorithms to solve the brought forward two models. Finally, extensive experiments on some real datasets are designed in the presence of reflection, large-scale rotation, stretch, noise, and outliers. Some state-of-the-art methods, including CPD, APM, RANSAC, TPS-ICP, TPS-RPM, RPMNet, and classical DOT methods, are to be discussed and compared. For different levels of degradation, the numerical results demonstrate that the proposed methods perform more favorably and robustly than the other methods.
As artificial intelligence and large data develop, distributed optimization shows the great potential in the research of machine learning, particularly deep learning. As an important distributed optimization problem, the nonsmooth distributed optimization problem over an undirected multi-agent system with inequality and equality constraints frequently appears in deep learning. To deal with this optimization problem cooperatively, a novel neural network with lower dimension of solution space is presented. It is demonstrated that the state solution of proposed approach can enter the feasible region. Also, it can also prove that the state solution achieves consensus and finally converges to the optimal solution set. Moreover, the proposed approach here does not depend on the boundedness of the feasible region, which is a necessary assumption in some simplified neural network. Finally, some simulation results and a practical application are given to reveal the efficacy and practicability.
We study a class of constrained sparse optimization problems with cardinality penalty, where the feasible set is defined by box constraint, and the loss function is convex but not necessarily smooth. First, we propose an accelerated smoothing hard thresholding (ASHT) algorithm for solving such problems, which combines smoothing approximation, extrapolation technique and iterative hard thresholding method. The extrapolation coefficients can be chosen to satisfy sup _k β _k=1 . We discuss the convergence of ASHT algorithm with different extrapolation coefficients, and give a sufficient condition to ensure that any accumulation point of the iterates is a local minimizer of the original problem. For a class of special updating schemes on the extrapolation coefficients, we obtain that the iterates are convergent to a local minimizer of the problem, and the convergence rate is o(ln ^σ k/k) with σ∈ (1/2, 1] on the loss and objective function values. Second, we consider the case in which the loss function is Lipschitz continuously differentiable, and develop an accelerated hard thresholding (AHT) algorithm to solve it. We prove that the iterates of AHT algorithm converge to a local minimizer of the problem that satisfies a desirable lower bound property. Moreover, we show that the convergence rates of loss and objective function values are o(k^-2) . Finally, some numerical examples are presented to show the theoretical results.
We propose a smoothing accelerated proximal gradient (SAPG) method with fast convergence rate for finding a minimizer of a decomposable nonsmooth convex function over a closed convex set. The proposed algorithm combines the smoothing method with the proximal gradient algorithm with extrapolation k-1/k+α -1 and α > 3 . The updating rule of smoothing parameter μ _k is a smart scheme and guarantees the global convergence rate of o(ln ^σk/k) with σ∈ (1/2,1] on the objective function values. Moreover, we prove that the iterates sequence is convergent to an optimal solution of the problem. We then introduce an error term in the SAPG algorithm to get the inexact smoothing accelerated proximal gradient algorithm. And we obtain the same convergence results as the SAPG algorithm under the summability condition on the errors. Finally, numerical experiments show the effectiveness and efficiency of the proposed algorithm.
Sparse optimization involving the L0-norm function as the regularization in objective function has a wide application in many fields. In this paper, we propose a projected neural network modeled by a differential equation to solve a class of these optimization problems, in which the objective function is the sum of a nonsmooth convex loss function and the regularization defined by the L0-norm function. This optimization problem is not only nonconvex, but also discontinuous. To simplify the structure of the proposed network and let it own better convergence properties, we use the smoothing method, where the new constructed smoothing function for the regularization term plays a key role. We prove that the solution to the proposed network is globally existent and unique, and any accumulation point of it is a critical point of the continuous relaxation model. Except for a special case, which can be easily justified, any critical point is a local minimizer of the considered sparse optimization problem. It is an interesting thing that all critical points own a promising lower bound property, which is satisfied by all global minimizers of the considered problem, but is not by all local minimizers. Finally, we use some numerical experiments to illustrate the efficiency and good performance of the proposed method for solving this class of sparse optimization problems, which include the most widely used models in feature selection of classification learning.