Through a "partial strong convexity" lemma, this paper shows how bounds on subproblem objective value suboptimality can be used in inexact augmented Lagrangian methods and ADMM algorithms. The ADMM result uses a small but important refinement on a long-standing criterion for approximately solving subproblems. The results enable two new approaches to computing Lagrangian bounds on the optimal values of stochastic mixed-integer programming problems, with simpler convergence analysis than the prior state of the art. In each case, the subproblems are solved by variants of the classical Frank-Wolfe algorithm. However, as compared to prior methods of the same type, there is much more freedom in the choice of Frank-Wolfe variant.
This paper presents two new techniques relating to inexact solution of subproblems in augmented Lagrangian methods for convex programming. The first involves combining a relative error criterion for solution of the subproblems with over- or under-relaxation of the multiplier update step. In one interpretation of our proposed iterative scheme, a predetermined amount of relaxation effects the criterion for an acceptably accurate solution value. Alternatively, the amount of multiplier step relaxation can be adapted to the accuracy of the subproblem subject to a viability test employing the discriminant of a certain quadratic function. The second innovation involves solution of augmented Lagrangian subproblems for problems posed in standard Fenchel-Rockafellar form. We show that applying alternating minimization to this subproblem, as in the first two steps of the ADMM, is equivalent to executing the classical proximal gradient method on a dual formulation of the subproblem. By substituting more sophisticated variants of the proximal gradient method for the classical one, it is possible to construct new ADMM-like methods with better empirical performance than using ordinary alternating minimization within an inexact augmented Lagrangian framework. The paper concludes by describing some computational experiments exploring using these two innovations, both separately and jointly, to solve LASSO problems.
We present a new, stochastic variant of the projective splitting (PS) family of algorithms for inclusion problems involving the sum of any finite number of maximal monotone operators. This new variant uses a stochastic oracle to evaluate one of the operators, which is assumed to be Lipschitz continuous, and (deterministic) resolvents to process the remaining operators. Our proposal is the first version of PS with such stochastic capabilities. We envision the primary application being machine learning (ML) problems, with the method’s stochastic features facilitating “mini-batch” sampling of datasets. Since it uses a monotone operator formulation, the method can handle not only Lipschitz-smooth loss minimization, but also min–max and noncooperative game formulations, with better convergence properties than the gradient descent-ascent methods commonly applied in such settings. The proposed method can handle any number of constraints and nonsmooth regularizers via projection and proximal operators. We prove almost-sure convergence of the iterates to a solution and a convergence rate result for the expected residual, and close with numerical experiments on a distributionally robust sparse logistic regression problem.
In “Projective Hedging Algorithms for Multistage Stochastic Programming, Supporting Distributed and Asynchronous Implementation,” Eckstein, Watson, and Woodruff derive a new class of decomposition methods for convex multistage stochastic programs defined on finite but potentially large scenario trees. These methods resemble Rockafellar and Wets’ now-classical progressive hedging (PH) method but are based on a flexible projective operator-splitting scheme instead of the standard alternating direction method of multipliers (ADMM). The new algorithms only need to reoptimize subproblems for a subset of the scenarios at each iteration, instead of all of them, and are also amenable to a form of asynchronous implementation, without the algorithm randomization or small step-size requirements usually imposed in such contexts. In the online appendix, the authors demonstrate significant computational gains over PH, applying hundreds or thousands of processor cores to problem instances with up to a million scenarios.
We present a new, stochastic variant of the projective splitting (PS) family of algorithms for monotone inclusion problems. It can solve min-max and noncooperative game formulations arising in applications such as robust ML without the convergence issues associated with gradient descent-ascent, the current de facto standard approach in such situations. Our proposal is the first version of PS able to use stochastic (as opposed to deterministic) gradient oracles. It is also the first stochastic method that can solve min-max games while easily handling multiple constraints and nonsmooth regularizers via projection and proximal operators. We close with numerical experiments on a distributionally robust sparse logistic regression problem.
We propose a decomposition algorithm for multistage stochastic programming that resembles the progressive hedging method of Rockafellar and Wets, but is provably ca-pable of several forms of asynchronous operation. We derive the method from a class of projective operator splitting methods fairly recently proposed by Combettes and Eckstein, significantly expanding the known applications of those methods. Our derivation assures convergence for convex problems whose feasible set is compact, subject to some standard regularity conditions and a mild “fairness” condition on subproblem selection. The method’s convergence guarantees are deterministic and do not require randomization, in contrast to some other proposed asynchronous variations of progressive hedging. We describe a distributed implementation of the method within the mpi-sppy system and present the results of computational experiments on up to a million scenarios and using as many as 2,400 processor cores in an HPC (high-performance computing) environment. These experiments evaluate the performance of the method when operating in a “block asynchronous” manner, meaning that it still alternates between non-overlapping decomposition and coordination phases, but only a subset of the subproblems are solved during each decomposition phase. Using a particular “greedy” heuristic to select which subproblems to solve, our experiments show that this tactic can make significantly more efficient use of computational resources than the original form of progressive hedging.
ROL-PEBBL is a C++, MPI-based parallel code for mixed-integer PDE-constrained optimization (MIPDECO). In these problems we wish to optimize (control, design, etc.) physical systems, which must obey the laws of physics, when some of the decision variables must take integer values. ROL-PEBBL combines a code to efficiently search over integer choices (PEBBL = Parallel Enumeration Branch-and-Bound Library) and a code for efficient nonlinear optimization, including PDE-constrained optimization (ROL = Rapid Optimization Library). In this report, we summarize the design of ROL-PEBBL and initial applications/results. For an artificial source-inversion problem, finding sources of pollution on a grid from sparse samples, ROL-PEBBLs solution for the nest grid gave the best optimization guarantee for any general solver that gives both a solution and a quality guarantee.
A recent innovation in projective splitting algorithms for monotone operator inclusions has been the development of a procedure using two forward steps instead of the customary resolvent step for operators that are Lipschitz continuous. This paper shows that the Lipschitz assumption is unnecessary when the forward steps are performed in finite-dimensional spaces: a backtracking linesearch yields a convergent algorithm for operators that are merely continuous with full domain.
This short paper describes a simple subgradient-based techniques for deriving bounds on the optimal solution value when using the ADMM to solve convex optimization problems. The technique requires a bound on the magnitude of some optimal solution vector, but is otherwise completely general. Some computational examples using LASSO problems demonstrate that the technique can produce steadily converging bounds in situations in which standard Lagrangian bounds yield little or no useful information. A second set of experiments establishes a proof of concept indicating the potential practical usefulness of the bounding technique.
This paper derives new inexact variants of the Douglas-Rachford splitting method for maximal monotone operators and the alternating direction method of multipliers (ADMM) for convex optimization. The analysis is based on a new inexact version of the proximal point algorithm that includes both an inertial step and overrelaxation. We apply our new inexact ADMM method to LASSO and logistic regression problems and obtain somewhat better computational performance than earlier inexact ADMM methods.
This work is concerned with the classical problem of finding a zero of a sum of maximal monotone operators. For the projective splitting framework recently proposed by Combettes and Eckstein, we show how to replace the fundamental subproblem calculation using a backward step with one based on two forward steps. The resulting algorithms have the same kind of coordination procedure and can be implemented in the same block-iterative and highly flexible manner, but may perform backward steps on some operators and forward steps on others. Prior algorithms in the projective splitting family have used only backward steps. Forward steps can be used for any Lipschitz-continuous operators provided the stepsize is bounded by the inverse of the Lipschitz constant. If the Lipschitz constant is unknown, a simple backtracking linesearch procedure may be used. For affine operators, the stepsize can be chosen adaptively without knowledge of the Lipschitz constant and without any additional forward steps. We close the paper by empirically studying the performance of several kinds of splitting algorithms on a large-scale rare feature selection problem.
This work describes a new variant of projective splitting for solving maximal monotone inclusions and complicated convex optimization problems. In the new version, cocoercive operators can be processed with a single forward step per iteration. In the convex optimization context, cocoercivity is equivalent to Lipschitz differentiability. Prior forward-step versions of projective splitting did not fully exploit cocoercivity and required two forward steps per iteration for such operators. Our new single-forward-step method establishes a symmetry between projective splitting algorithms, the classical forward–backward splitting method (FB), and Tseng’s forward-backward-forward method. The new procedure allows for larger stepsizes for cocoercive operators: the stepsize bound is $$2\beta$$ for a $$\beta$$ -cocoercive operator, the same bound as has been established for FB. We show that FB corresponds to an unattainable boundary case of the parameters in the new procedure. Unlike FB, the new method allows for a backtracking procedure when the cocoercivity constant is unknown. Proving convergence of the algorithm requires some departures from the prior proof framework for projective splitting. We close with some computational tests establishing competitive performance for the method.
This article describes a new rule-enhanced penalized regression procedure for the generalized regression problem of predicting scalar responses from observation vectors in the absence of a preferred functional form. It enhances standard L1-penalized regression by adding dynamically generated rules, that is, new 0-1 covariates, corresponding to multidimensional “box” sets. In contrast to prior approaches to this class of problems, we draw heavily on standard (but non-polynomial-time) mathematical programming techniques, enhanced by parallel computing. We identify and incorporate new rules using a form of classical column generation and solve the resulting pricing subproblem, which is NP-hard, either exactly by a specialized parallel branch-and-bound method or by a greedy heuristic based on Kadane’s algorithm. The resulting rule-enhanced regression method can be computation intensive when we solve the subproblems exactly, but our computational tests suggest that it outperforms prior methods at making accurate and stable predictions from relatively small data samples. Through selective use of our greedy heuristic, we can make our method’s run time generally competitive with some established methods, without sacrificing prediction performance. We call our method’s pricing subproblem rectangular maximum agreement.
OF THE DISSERTATION Risk-averse Decision Making and Bilevel Stochastic Programming with Applications By Deniz Seyed Eskandani Dissertation Director: Jonathan Eckstein We present a novel modeling approach to time-consistently formulate three-stage riskaverse stochastic programming problems, using bilevel programming. For certain classes of applications, we empirically demonstrate that our approach can behave substantially differently from prior formulations of the problem. To obtain these results, we reformulate the NP-hard bilevel model using complementarity constraints and then express it as a disjunctive program. However, this approach does not scale well, even using the best available commercial MIP solvers. To overcome this hurdle, we use a proximal bundle method to efficiently find a lower bound for the optimal solution. We further supplement this procedure with an upper bound by proposing an approach to find a feasible solution. We implement our algorithm in the gurobipy module of Python and apply it to various classes of problems and compare our computational results with our earlier disjunctive programming approach. We find that our bounds can provide a better approximation of the optimal solution than the MIP-solver approach and can scale to larger problems.
Projective splitting is a family of methods for solving inclusions involving sums of maximal monotone operators. First introduced by Eckstein and Svaiter in 2008, these methods have enjoyed significant innovation in recent years, becoming one of the most flexible operator-splitting frameworks available. While weak convergence of the iterates to a solution has been established, there have been few attempts to study convergence rates of projective splitting. The aim of this paper is to do so under various assumptions. To this end, it makes four main contributions. First, in the context of convex optimization, an $O(1/k)$ ergodic function convergence rate is established. Second, for strongly monotone inclusions, strong convergence is established as well as an ergodic $O(1/\sqrt{k})$ convergence rate for the distance from the iterates to the solution. Third, for inclusions featuring strong monotonicity and cocoercivity, linear convergence is established. We also consider the special case of one operator. In this case we show that projective splitting reduces to either the extragradient method or the proximal-point method, depending on whether forward or backward steps are used.
This paper presents an efficient technique for matrix-vector and vector-transpose-matrix multiplication in distributed-memory parallel computing environments, where the matrices are unstructured, sparse, and have a substantially larger number of columns than rows or vice versa. Our method allows for parallel I/O, does not require extensive preprocessing, and has the same communication complexity as matrix-vector multiplies with column or row partitioning. Our implementation of the method uses MPI. We partition the matrix by individual nonzero elements, rather than by row or column, and use an "overlapped" vector representation that is matched to the matrix. The transpose multiplies use matrix-specific MPI communicators and reductions that we show can be set up in an efficient manner. The proposed technique achieves a good work per processor balance even if some of the columns are dense, while keeping communication costs relatively low.
We propose new primal-dual decomposition algorithms for solving systems of inclusions involving sums of linearly composed maximally monotone operators. The principal innovation in these algorithms is that they are block-iterative in the sense that, at each iteration, only a subset of the monotone operators needs to be processed, as opposed to all operators as in established methods. Flexible strategies are used to select the blocks of operators activated at each iteration. In addition, we allow lags in operator processing, permitting asynchronous implementation. The decomposition phase of each iteration of our methods is to generate points in the graphs of the selected monotone operators, in order to construct a half-space containing the Kuhn–Tucker set associated with the system. The coordination phase of each iteration involves a projection onto this half-space. We present two related methods: the first method provides weakly convergent primal and dual sequences under general conditions, while the second is a variant in which strong convergence is guaranteed without additional assumptions. Neither algorithm requires prior knowledge of bounds on the linear operators involved or the inversion of linear operators. Our algorithmic framework unifies and significantly extends the approaches taken in earlier work on primal-dual projective splitting methods.
We derive a new approximate version of the alternating direction method of multipliers (ADMM) which uses a relative error criterion. The new version is somewhat restrictive and allows only one of the two subproblems to be minimized approximately, but nevertheless covers commonly encountered special cases. The derivation exploits the long-established relationship between the ADMM and both the proximal point algorithm (PPA) and Douglas–Rachford (DR) splitting for maximal monotone operators, along with a relative-error of the PPA due to Solodov and Svaiter. In the course of analysis, we also derive a version of DR splitting in which one operator may be evaluated approximately using a relative error criterion. We computationally evaluate our method on several classes of test problems and find that it significantly outperforms several alternatives on one problem class.
This paper presents two new approximate versions of the alternating direction method of multipliers (ADMM) derived by modifying of the original "Lagrangian splitting" convergence analysis of Fortin and Glowinski. They require neither strong convexity of the objective function nor any restrictions on the coupling matrix. The first method uses an absolutely summable error criterion and resembles methods that may readily be derived from earlier work on the relationship between the ADMM and the proximal point method, but without any need for restrictive assumptions to make it practically implementable. It permits both subproblems to be solved inexactly. The second method uses a relative error criterion and the same kind of auxiliary iterate sequence that has recently been proposed to enable relative-error approximate implementation of non-decomposition augmented Lagrangian algorithms. It also allows both subproblems to be solved inexactly, although ruling out "jamming" behavior requires a somewhat complicated implementation. The convergence analyses of the two methods share extensive underlying elements.
Donald Goldfarb合作论文数Department of Industrial Engineering and Operations Research, Columbia University1