
In this work, we propose a flexible block coordinate method for unconstrained optimization problems under Hölder continuity assumptions. The method guarantees convergence to stationary points and has worst-case complexity results comparable to those obtained by single-block methods that assume Lipschitz or Hölder continuity. The approach is based on quadratic models of the objective function combined with quadratic regularization. Flexible block selection strategies are allowed, and different sufficient descent conditions previously considered in the literature are unified within the same framework. The proposed method is fully implementable and incorporates a practical certificate of stationarity as a stopping criterion. Well-definiteness and convergence are established under the Hölder continuity of the objective function gradient, extending classical results that typically rely on Lipschitz assumptions. Illustrative numerical tests are performed and a freely available Julia implementation is provided.
We consider a production plant consuming electric energy for its manufacturing activities. We assume that the plant is partially powered by on-site generated renewable energy and supported by an energy storage system. We seek to simultaneously plan the production activities and the energy supply of this plant. Uncertainty in renewable energy availability introduces a significant challenge to cost-effective decision-making. We develop a two-stage stochastic programming methodology to address the uncertainties associated with local renewable energy production. This modeling approach results in the formulation of a large-size mixed-integer linear program displaying a block-decomposable structure. To solve this problem, we introduce an enhanced branch-and-Benders-cut algorithm which combines several strategies to accelerate convergence towards an optimal solution. Computational results carried out on randomly generated medium to large-sized instances are presented. They show that the enhanced branch-and-Benders-cut algorithm clearly outperforms both a cut-and-branch algorithm and a standard branch-and-Benders-cut algorithm at solving the problem, especially for instances involving a large number of scenarios.
Neural differential-algebraic systems of equations (DAEs) are a modeling paradigm where some unknown relationships within a DAE are modeled with a neural network and learned from data. Training neural DAEs is more challenging than training neural ordinary-differential equations (ODEs), particularly for higher-index systems. Existing approaches utilize differentiable pipelines, usually comprising integration, projection, operator splitting, or penalty terms for algebraic constraints. The parameters are then updated in a sequential manner using gradient descent. In this work, we employ the simultaneous approach for DAE-constrained parameter estimation instead. This defines a fully discretized nonlinear programing problem (NLP), whose solution simultaneously obtains the neural network parameters and the trajectories of the corresponding DAE, while enforcing constraint satisfaction at the discretization points. We show that with careful initialization and handling of the neural network terms, this approach can be efficient for smaller-scale problems, including higher-index DAEs. As the number of parameters or the amount of data increase, decomposition strategies are necessary to make this method scalable. We present a bi-level approach, where a tractable, discretized NLP is solved at every iteration of an outer gradient descent method, and the sensitivity of its solution with respect to the neural network parameters is evaluated in an efficient manner. We demonstrate the scalability of this decomposition with respect to the number of parameters and the size of the training data set.
Recent work has shown that, in several growth settings, the Frank–Wolfe algorithm (FW) with fixed-parameter open-loop step-sizes η _t = ℓ/t+ℓ for ℓ∈ℕ_≥ 2 can converge faster than the classical 𝒪(t^-1) rate. In the strong-growth regime, increasing ℓ improves the asymptotic rate, so no single fixed choice of ℓ is uniformly preferable. We study the broader family of adaptive open-loop step-sizes η _t = g(t)/t+g(t) and extend the affine-invariant analysis from the fixed-parameter case to non-decreasing functions g for which the resulting step-size sequence is non-increasing. For the log-adaptive choice g(t)=2+log (t+1) , this yields a schedule that matches the asymptotic rates of fixed-parameter rules up to polylogarithmic factors in the weak-growth settings considered here, while improving the asymptotic sublinear rate in strong-growth settings. We also report numerical experiments demonstrating the efficacy of the new proposed step-size rule.
In this study, we propose a new exact approach to the competitive facility location problem with a limited choice rule. The problem consists in locating facilities to capture a market share that maximizes the profits of a new company entering into a competitive marketplace, knowing that customers patronize a limited number of the nearest facilities, i.e., they have a limited choice of facilities to favor. We propose an outer approximation approach devised to handle the nonlinear nature of the problem, with its master problem being solved using a multicut-variant Benders’ decomposition algorithm. Given the designed decomposition structure of the outer approximation master problem, we can analytically separate Benders’ feasibility cuts for each customer at a time. Our outer approximation-Benders’ decomposition algorithm outperformed existing methods in the literature when solving medium- and large-scale instances, being, on average, seven and two times faster, respectively, than the fastest previously reported method.
Dealing with uncertainty in optimization parameters is an important and longstanding challenge. Typically, uncertain parameters are predicted accurately, and then a deterministic optimization problem is solved. However, the decisions produced by this so-called predict-then-optimize procedure can be highly sensitive to inaccuracies. In this work, we contribute to recent efforts in producing decision-focused predictions, i.e., to build predictive models that are constructed with the goal of minimizing a regret measure on the decisions taken with them. We begin by formulating the exact expected regret minimization as a pessimistic bilevel optimization model. Then, we show computational complexity results of this problem for the case when the predictive model is a linear regression. This includes its membership in NP, which, in combination with a known NP-hardness result, establishes NP-completeness. Using duality arguments, we reformulate the pessimistic formulation exactly as a single-level problem; when the model is a linear regression, the resulting problem is a non-convex quadratically constrained quadratic optimization problem. Finally, leveraging the quadratic reformulation, we show various computational techniques to achieve empirical tractability. We report extensive computational results on shortest-path and bipartite matching instances with uncertain cost vectors. Our results indicate that our approach can work well in combination with many state-of-the-art decision-focused learning methods, improving their training and test performance.
This paper introduces a nonmonotone extension of the Front Descent method, a state-of-the-art descent-based algorithmic framework designed to approximate the Pareto front of smooth multiobjective optimization problems. The proposed approach incorporates novel nonmonotone line search strategies that allow temporary increases in some objective functions, potentially accelerating convergence and improving overall efficiency. The theoretical analysis demonstrates that the sequence of sets generated by the algorithm retains the convergence properties of the original framework. Establishing these properties in the nonmonotone setting is nontrivial: the loss of monotonicity across the set sequence introduces substantial analytical challenges, necessitating a careful and rigorous adaptation of the original convergence arguments. Finally, numerical results in the bound-constrained setting are presented to validate the goodness of the proposed approach.
Many optimization problems in machine learning can be studied through the lens of Riemannian nonsmooth optimization, where a nonsmooth objective is originally constrained to lie on a Riemannian manifold. In addition to this geometric view, the multiobjective approach has expansive applications in such problems, as one may aim to optimize more than one objective simultaneously, which makes evaluating trade-offs between objectives essential. In this paper, we study a retraction-based trust-region algorithm for solving nonsmooth multiobjective optimization problems defined on Riemannian manifolds. Considering general Riemannian manifolds, we analyze the convergence behavior of the proposed method. This general approach enables one to apply the proposed algorithm to a wider variety of cases, from unconstrained problems on Euclidean spaces to problems with manifold constraints. Furthermore, we study the application of the method in matrix machine learning problems such as fair sparse PCA problems, and multiobjective ℓ _1 -regularized least squares problems. We demonstrate the efficiency of the algorithm by reporting the numerical experiments on the mentioned applications.
The paper is devoted to the study of the Newton-type scheme applied to the problem of finding a solution to the inclusion governed by set-valued mappings. We develop a stability result for global metric regularity and use it to investigate the semilocal convergence for the Josephy-Newton framework. On the other hand, we investigate in the second fold of the paper the behavior of iterative process based on Josephy-Newton scheme for solving variational problems involving convex processes.
This paper presents a multi-faceted approach to solving the quadratic multidimensional knapsack problem (QMDKP), an NP-hard nonlinear combinatorial optimization problem that has received limited attention in the literature. We introduce a new testbed of QMDKP instances with induced profit-weight correlations and evaluate four linearization strategies for solving the problem exactly. To address scalability, we propose several heuristic and metaheuristic methods, including a greedy algorithm, a supervised machine learning (ML)-based approach, and a genetic algorithm. Our underlying ML framework employs binary classification to predict item inclusion probabilities, guiding a greedy construction process. Building on this framework, we explore hybrid strategies such as reduce-fill, expand-repair, and a genetic algorithm with ML-driven repair operators. Extensive computational results show that these ML-enhanced hybrid methods improve upon their baseline heuristics and that the genetic algorithm consistently yields near-optimal solutions. This work provides a generalizable template for hybrid optimization and learning-based approaches to challenging combinatorial problems.
This paper focuses on mathematical modeling of shape design, shape sensitivity analysis, and numerical implementation of shape optimization constrained by the time-dependent Stokes flow. The optimization models aim to minimize two classical shape functionals: energy dissipation and least-squares for velocity tracking. A porous medium model is applied to formulate the time-dependent Stokes equations on a fixed domain consisting of both fluid and solid regions. The existence of a solution to the perimeter regularized shape optimization problem is proved. For shape sensitivity analysis, the existence of the material derivative of the PDE solution and the Eulerian derivative of the shape functional have been rigorously established. Convected level set evolution governed by the distributed Eulerian derivative is applied to address shape design problems, such as pipe design, obstacle shape design, and shape reconstruction. Numerical examples demonstrate the effectiveness of the proposed algorithms in various scenarios.
In this paper, we revisit the convergence theory of the inexact restoration paradigm for nonlinear optimization. The paper first identifies the basic elements of a globally convergent method based on merit functions. Then, the inexact restoration method that employs a two-phase iteration is introduced as a special case. A specific implementation is presented that is supported in the solution of regularized subproblems. The proposed inexact restoration method includes more freedom in the computation of iterates and a novel procedure that integrates the computation of the penalty parameter with the optimization phase. It also includes an acceleration step based on solving a classic quadratic programming subproblem. The proposed theoretical framework is broad and flexible enough to provide a convergence theory for inexact restoration applications tailored to specific problems. Theoretical results include asymptotic convergence theory as well as a worst-case iteration and evaluation complexity analysis of the introduced methods. The paper concludes with numerical experiments that illustrate the practical performance of two variations of the proposed method for solving general nonlinear programming problems.
As inherently transparent models, classification trees play a central role in interpretable machine learning by providing easily traceable decision paths that allow users to understand how input features contribute to specific predictions. In this work, we introduce a new class of interpretable binary classification models, named Pareto-optimal trees, which aim at combining the complementary strengths of Optimal Classification Trees with Hyperplane-based splits (OCT-H) and Support Vector Machines (SVM). We formulate a bi-objective mixed-integer quadratic optimization problem, whose nondominated solutions represent trade-offs between these two different classification techniques. To further enhance robustness and performance, we propose the Pareto forest, an ensemble method based on the Pareto-optimal trees, aggregated through majority voting. Extensive experiments on benchmark datasets demonstrate that our models can outperform standard methods such as CART and OCT, underscoring the improvements gained through the bi-objective perspective. In particular, Pareto-optimal trees unify the ability of OCT-H and SVM within a single framework, resulting in enhanced classification performance relative to either method alone. Embracing a multiobjective perspective allows the construction of multiple “high quality" trees. Our comparison between Pareto forests and random forests shows that building shallow ensembles from a small number of such optimized trees outperforms relying on a large set of random trees with variable depth.
We derive, from unified principles, local convergence and rate-of-convergence results for the classical Gauss-Newton method in a variety of settings. These include overdetermined and underdetermined systems of equations, constrained and unconstrained, possibly with inexact solution of subproblems, as well as the projected variant in the constrained case. Moreover, by a counter-example we show that contrary to some results claimed in the literature, the projected Gauss-Newton method in general does not converge superlinearly under any reasonable assumptions. We then establish its linear rate of convergence.
In the context of fault-detection problems, the objective is to identify all defective items among a set of n binary-state items using the minimum number of tests. The group testing paradigm, which allows testing a subset of items in a single test, serves as a fundamental technique for efficiently classifying large populations. We study a central problem in the combinatorial group testing model where the number d of defective items is unknown in advance. Let M_α (d|n) denote the maximum number of tests required by an algorithm α for this problem, and M(d, n) denote the minimum number of tests required in the worst case when d is known in advance. An algorithm α is called a c-competitive algorithm if there exist constants c and a such that, for 0≤ d < n , M_α(d|n)≤ cM(d,n)+a . We develop a modular two-stage framework for this problem: (i) a constant-size preliminary-testing stage that extracts coarse information about d; and (ii) a conditional invocation stage that, based on the preliminary outcome, invokes either our newly developed up-zig-zag approach or an existing strongly competitive algorithm. With carefully designed switching rules, this framework yields a deterministic adaptive algorithm with a competitive ratio c ≤ 1.431 , improving the previous best-known bound of 1.452. This guarantee is achieved via a reusable integration principle that combines a constant preliminary overhead with a second-stage choice tailored to the remaining uncertainty.
Multiple projects are generally executed concurrently in an organization. This study discusses selecting an optimal subset of projects from a set of candidate projects to maximize benefits under limited resources. The mathematical model considers project interdependency and cardinality constraints. It is challenging to solve for a globally optimal solution. Heuristic methods cannot guarantee the global optimality of the obtained solution, and exact methods with complicated model transformations are inefficient. This study utilizes a novel linearization approach to efficiently transform the project portfolio selection problem with pairwise and three-way cross-product terms as a mixed-integer linear program. Compared with current transformation methods, the proposed method reduces numerous continuous variables required to linearize the original problem and significantly enhances the computational efficiency. Additionally, the effectiveness and practicability of the proposed method are demonstrated using data from a semiconductor company. The principles of project portfolio selection, such as optimizing resource allocation, managing interdependencies, and maximizing benefits, are applicable across various industries. The proposed method can help organizations in different sectors efficiently select an optimal project portfolio under limited resources.
A scalable modulus-based matrix splitting (SMMS) method is extended for solving the vertical nonlinear complementarity problem (VNCP) which shuns any auxiliary variables. With the aim to further enhance the efficiency, by introducing the relaxation matrix based on SMMS method, the relaxed scalable modulus-based matrix splitting (RSMMS) method is presented. When s=2 , a comparison theorem between SMMS and RSMMS methods is presented to theoretically demonstrate the effectiveness of introducing the relaxation matrix. Under mild conditions, the convergence of the RSMMS method for arbitrary s is established and the RSMMS method is proved to converge to the unique solution of the VNCP. As a by-product, sufficient conditions are established for a VNCP to have a unique solution. Numerical results are given to demonstrate the efficiency of SMMS and RSMMS methods.
This paper proposes a multilevel hypergraph partitioning framework that integrates both graph and hypergraph structures. In the multilevel solution process, multiple candidate partitions are generated in parallel during the initial partitioning stage and independently refined, with the best solution ultimately selected. This strategy effectively reduces the risk of the algorithm being trapped in local optima and improves solution stability and robustness. For the initial partitioning, the discrete hypergraph problem is transformed into a continuous unconstrained optimization model, which is efficiently solved by the existing solver CGOPT 2.0 to obtain high-quality low-dimensional embeddings. The continuous solutions are then mapped to 2-way partitions via hyperplane rounding, while general k-way partitions are constructed recursively, maintaining both partition quality and scalability. On the theoretical side, leveraging CGOPT 2.0, we prove that the objective function of the constructed continuous model satisfies the KŁ property and further derive explicit convergence rates for both the function value and iteration sequences, providing rigorous guarantees of reliability and numerical stability in nonconvex settings. Comprehensive experiments on public benchmarks such as ISPD98 and Titan23 demonstrate that our algorithm generally achieves superior cutsize quality compared to KaHyPar, hMetis, K-SpecPart, Mt-KaHyPar and PaToH, with comparable or slightly higher runtime, and ablation studies further validate the effectiveness of each novel component.
The p-hub median problem is central to location science, with multiple applications in network design. Given a set of nodes, each of them sends and receives some flow to and from the other nodes. Among these nodes, p have to be selected to serve as hubs, which act as switching points for directing the flow. These hubs are interconnected with direct links to form a complete subnetwork, while every non-hub node is linked to one of these selected hubs. The p-hub median problem typically arises in the modeling of air traffic, disaster management, logistics, and telecommunications networks, among others. In a world driven by dynamic data, situations may arise in which the same p-hub median problem must be solved for multiple instances with varying flows. While solving each instance exactly from scratch may require long computational times, learning from a training set of previously solved instances can help predict near-optimal solutions for new ones in which the flows have been changed. In this setting, we propose two heuristics to address the p-hub median location problem with single allocation and unconstrained hub capacity: one for problems in which all hub links are activated, allowing flow exchange between every pair of hubs, and another for problems in which flow exchanges between hubs are limited because not all hub links are activated. For the case where all hubs exchange flow with each other, the set of nodes acting as hubs is predicted using a machine learning approach. In the case where flow exchanges between hubs are limited, both the hubs and the activated connections between hubs are predicted using machine learning. Both heuristics include carefully designed strategies to restore feasibility. Our extensive computational results demonstrate the effectiveness of the proposed heuristics, both in terms of the good quality of the solutions, which are near-optimal, and in their ability to produce them in record time when compared to state-of-the-art approaches.
Stochastic multi-objective optimization (SMOO) has attracted increasing interest as a framework for machine learning problems with multiple objectives. The nonlinearity of the subproblem solution mapping introduces bias in the weighting parameters, which complicates the convergence analysis of multi-gradient algorithms. In this paper, we propose the Multi-gradient Stochastic Mirror Descent (MSMD) algorithm, which solves the SMOO subproblem via the stochastic mirror descent method and provides convergence guarantees. By selecting an appropriate Bregman distance-generating function, the algorithm admits a closed-form solution for the weighting vector and requires only a single stochastic gradient sample per iteration. We establish sublinear convergence rates for MSMD under four combinations of inner and outer step size strategies. A variant of MSMD for SMOO with preferences is also proposed and analyzed. Numerical experiments on benchmark test functions and neural network training tasks show that MSMD achieves competitive performance.