High-order tensor methods that employ local Taylor models of degree p within adaptive regularization frameworks (ARp) have recently received significant attention, due to their improved/optimal global and local rates of convergence, for both convex and nonconvex optimization problems. The numerical performance of tensor methods for general unconstrained optimization problems remains insufficiently explored/understood, which we address in this paper, by showcasing the numerical performance of standard second- and third-order variants ( p=2,3 ) and proposing novel techniques for key algorithmic aspects when p≥ 3 , to improve the numerical efficiency of tensor variants. In particular, to improve the adaptive choice of the regularization parameter, we extend the interpolation-based updating strategy introduced in [Gould, Porcelli and Toint, Comput Optim Appl (2012) 53:1–22] for p=2 , to the case when p ≥ 3 . We identify fundamental differences between the different local minima of the regularised subproblems for p=2 and p ≥ 3 and their effect on algorithm performance. Then, when p≥ 3 , we introduce a novel pre-rejection technique that rejects poor/unsuccessful subproblem minimizers (that we refer to as ’transient’) prior to any function evaluation, thereby reducing cost and selecting useful (’persistent’) ones. Numerical studies showcase the efficiency improvements generated by our proposed modifications of the AR3 algorithm. We also assess numerically, the effect of different subproblem termination conditions and choice of initial regularization parameter on the overall algorithm performance. Finally, we benchmark our best-performing AR3 variants, as well as those in [Birgin et al, Optim Lett (2020) 14:815–838], against second-order ones (AR2). Encouraging results on standard test problems are obtained, confirming that AR3 variants can be made to outperform second-order variants in terms of objective evaluations, derivative evaluations, and number of subproblem solves. We provide an efficient, extensive and modular software package in MATLAB that includes many AR2 and AR3 variants, including Hessian- and tensor-free ones, allowing ease of use and experimentation for interested users.
For the case of inflexible demand and considering network constraints, we introduce a Cost Minimisation (CM) based market clearing mechanism, and a model representing the standard Social Welfare Maximisation mechanism used in European Day Ahead Electricity Markets. Since the CM model corresponds to a more challenging optimisation problem, we propose four numerical algorithms that leverage the problem structure, each with different trade-offs between computational cost and convergence guarantees. These algorithms are evaluated on synthetic data to provide some intuition of their performance. We also provide strong (but partial) analytical results to facilitate efficient solution of the CM problem, which call for the introduction of a new concept: optimal zonal stack curves, and these results are used to devise one of the four solution algorithms. An evaluation of the CM and SWM models and their comparison is performed, under the assumption of truthful bidding, on the real world data of Central Western European Day Ahead Power Market during the period of 2019-2020. We show that the SWM model we introduce gives a good representation of the historical time series of the real prices. Further, the CM reduces the market power of producers, as generally this results in decreased zonal prices and always decreases the total cost of electricity procurement when compared to the currently employed SWM.
Quasi-Newton methods employ an update rule that gradually improves the Hessian approximation using the already available gradient evaluations. We propose higher-order secant updates which generalize this idea to higher-order derivatives, approximating, for example, third derivatives (which are tensors) from given Hessian evaluations. Our generalization is based on the observation that quasi-Newton updates are least-change updates satisfying the secant equation, with different methods using different norms to measure the size of the change. We present a full characterization for least-change updates in weighted Frobenius norms (satisfying an analogue of the secant equation) for derivatives of arbitrary order. Moreover, we establish convergence of the approximations to the true derivative under standard assumptions and explore the quality of the generated approximations in numerical experiments.
Binary matrix factorization is an essential tool for identifying discrete patterns in binary data. In this paper, we consider the rank-k binary matrix factorization problem (k-BMF) under Boolean arithmetic: we are given an n × m binary matrix X with possibly missing entries and need to find two binary matrices A and B of dimension n × k and k × m, respectively, which minimize the distance between X and the Boolean product of A and B in the squared Frobenius distance. We present a compact and two exponential size integer programs (IPs) for k-BMF and show that the compact IP has a weak linear programming (LP) relaxation, whereas the exponential size IPs have a stronger equivalent LP relaxation. We introduce a new objective function, which differs from the traditional squared Frobenius objective in attributing a weight to zero entries of the input matrix that is proportional to the number of times the zero is erroneously covered in a rank-k factorization. For one of the exponential size Ips, we describe a computational approach based on column generation. Experimental results on synthetic and real-world data sets suggest that our integer programming approach is competitive against available methods for k-BMF and provides accurate low-error factorizations. Funding: O. Günlük was partially supported by the Office of Naval Research [Grant N00014-21-1-2575]. R. Á. Kovács was supported by a doctoral scholarship from The Alan Turing Institute under the EPSRC [Grant EP/N510129/1] and the Office for National Statistics. R. A. Hauser was supported by The Alan Turing Institute under the EPSRC [Grant EP/N510129/1].
In the seminal paper on optimal execution of portfolio transactions, Almgren and Chriss (2001) define the optimal trading strategy to liquidate a fixed volume of a single security under price uncertainty. Yet there exist situations, such as in the power market, in which the volume to be traded can only be estimated and becomes more accurate when approaching a specified delivery time. During the course of execution, a trader should then constantly adapt their trading strategy to meet their fluctuating volume target. In this paper, we develop a model that accounts for volume uncertainty and we show that a risk-averse trader has benefit in delaying their trades. More precisely, we argue that the optimal strategy is a trade-off between early and late trades in order to balance risk associated with both price and volume. By incorporating a risk term related to the volume to trade, the static optimal strategies suggested by our model avoid the explosion in the algorithmic complexity usually associated with dynamic programming solutions, all the while yielding competitive performance.
Network constraints play a key role in the price finding mechanism for European Power Markets, but historical data is very sparse and usually insufficient for many quantitative applications. We reconstruct the constraints data, known as the Power Transmission Distribution Factors (PTDFs) and Remaining Available Margins (RAMs), by first recovering the underlying time dependent signals known as the Generation Shift Keys (GSKs) and Phase Angles (PAs), and the electricity grid characteristics, via a mathematical optimisation problem. This is solved by exploiting marginal convexity in certain subspaces via alternating minimisation. The GSKs and PAs are then mapped to the PTDFs and RAMs, using the grid structure. Our reconstruction achieves in-sample relative errors of 22.3% and 15.9% for the PTDFs and RAMs respectively, while the out-of-sample ones are 25.4% and 18.1% respectively. We further show that our model outperforms the naive approach, and that the reconstructed GSKs and PAs recover specific structure.
Spot electricity markets are considered under a Game-Theoretic framework, where risk averse players submit orders to the market clearing mechanism to maximise their own utility. Consistent with the current practice in Europe, the market clearing mechanism is modelled as a Social Welfare Maximisation problem, with zonal pricing, and we consider inflexible demand, physical constraints of the electricity grid, and capacity-constrained producers. A novel type of non-parametric risk aversion based on a defined worst case scenario is introduced, and this reduces the dimensionality of the strategy variables and ensures boundedness of prices. By leveraging these properties we devise Jacobi and Gauss-Seidel iterative schemes for computation of approximate global Nash Equilibria, which are in contrast to derivative based local equilibria. Our methodology is applied to the real world data of Central Western European (CWE) Spot Market during the 2019-2020 period, and offers a good representation of the historical time series of prices. By also solving for the assumption of truthful bidding, we devise a simple method based on hypothesis testing to infer if and when producers are bidding strategically (instead of truthfully), and we find evidence suggesting that strategic bidding may be fairly pronounced in the CWE region.
Identifying discrete patterns in binary data is an important dimensionality reduction tool in machine learning and data mining. In this paper, we consider the problem of low-rank binary matrix factorisation (BMF) under Boolean arithmetic. Due to the hardness of this problem, most previous attempts rely on heuristic techniques. We formulate the problem as a mixed integer linear program and use a large scale optimisation technique of column generation to solve it without the need of heuristic pattern mining. Our approach focuses on accuracy and on the provision of optimality guarantees. Experimental results on real world datasets demonstrate that our proposed method is effective at producing highly accurate factorisations and improves on the previously available best known results for 15 out of 24 problem instances.
Principal component analysis (PCA) is an important pattern recognition and dimensionality reduction tool in many applications. Principal components are computed as eigenvectors of a maximum likelihood covariance $\widehat{\varSigma }$ that approximates a population covariance $\varSigma$, and these eigenvectors are often used to extract structural information about the variables (or attributes) of the studied population. Since PCA is based on the eigendecomposition of the proxy covariance $\widehat{\varSigma }$ rather than the ground-truth $\varSigma$, it is important to understand the approximation error in each individual eigenvector as a function of the number of available samples. The combination of recent results of Koltchinskii & Lounici (2017, Bernoulli, 23, 110–133) and Yu et al. (2015, Biometrika, 102, 315–323) yields such bounds. In the present paper we sharpen these bounds and show that eigenvectors can often be reconstructed to a required accuracy from a sample of strictly smaller size order.
In this paper, we construct an algorithm for minimising piecewise smooth functions for which derivative information is not available. The algorithm constructs a pair of quadratic functions, one on each side of the point with smallest known function value, and selects the intersection of these quadratics as the next test point. This algorithm relies on the quadratic function underestimating the true function within a specific range, which is accomplished using a adjustment term that is modified as the algorithm progresses.
Considering optimal alignments of two i.i.d. random sequences of length n, we show that for Lebesgue-almost all scoring functions, almost surely the empirical distribution of aligned letter pairs in all optimal alignments converges to a unique limiting distribution as n tends to infinity. This result helps understanding the microscopic path structure of a special type of last-passage percolation problem with correlated weights, an area of long-standing open problems. Characterizing the microscopic path structure also yields robust alternatives to the use of optimal alignment scores alone for testing the homology of genetic sequences.
Principal component analysis (PCA) finds the best linear representation of data and is an indispensable tool in many learning and inference tasks. Classically, principal components of a dataset are interpreted as the directions that preserve most of its "energy," an interpretation that is theoretically underpinned by the celebrated Eckart-Young-Mirsky theorem. This paper introduces many other ways of performing PCA, with various geometric interpretations, and proves that the corresponding family of nonconvex programs has no spurious local optima, while possessing only strict saddle points. These programs therefore loosely behave like convex problems and can be efficiently solved to global optimality, for example, with certain variants of the stochastic gradient descent. Beyond providing new geometric interpretations and enhancing our theoretical understanding of PCA, our findings might pave the way for entirely new approaches to structured dimensionality reduction, such as sparse PCA and nonnegative matrix factorization. More specifically, we study an unconstrained formulation of PCA using determinant optimization that might provide an elegant alternative to the deflating scheme commonly used in sparse PCA.
This paper introduces Memory-limited Online Subspace Estimation Scheme (MOSES) for both estimating the principal components of streaming data and reducing its dimension. More specifically, in various applications such as sensor networks, the data vectors are presented sequentially to a user who has limited storage and processing time available. Applied to such problems, MOSES can provide a running estimate of leading principal components of the data that has arrived so far and also reduce its dimension. MOSES generalises the popular incremental Singular Vale Decomposition (iSVD) to handle thin blocks of data, rather than just vectors. This minor generalisation in part allows us to complement MOSES with a comprehensive statistical analysis, thus providing the first theoretically-sound variant of iSVD, which has been lacking despite the empirical success of this method. This generalisation also enables us to concretely interpret MOSES as an approximate solver for the underlying non-convex optimisation program. We find that MOSES consistently surpasses the state of the art in our numerical experiments with both synthetic and real-world datasets, while being computationally inexpensive.
Compressed sensing is a powerful mathematical modelling tool to recover sparse signals from undersampled measurements in many applications, including medical imaging. A large body of work investigates the case with linear measurements, while compressed sensing with nonlinear measurements has been considered more recently. We continue this line of investigation by considering a novel type of nonlinearity with special structure that occurs in data acquired by multi-emitter X-ray tomosynthesis systems with spatio-temporal overlap. In [15] we proposed a nonlinear optimization model to deconvolve the overlapping measurements. In this paper we propose a model that exploits the structure of the nonlinearity and a nonlinear tomosynthesis algorithm that has a practical running time of solving only two linear subproblems at the equivalent resolution. We underpin and justify the algorithm by deriving RIP bounds for the linear subproblems and conclude with numerical experiments that validate the approach.
Principal Component Analysis (PCA) finds the best linear representation for data and is an indispensable tool in many learning tasks. Classically, principal components of a dataset are interpreted as the directions that preserve most of its "energy", an interpretation that is theoretically underpinned by the celebrated Eckart-Young-Mirsky Theorem. There are yet other ways of interpreting PCA that are rarely exploited in practice, largely because it is not known how to reliably solve the corresponding non-convex optimisation programs. In this paper, we consider one such interpretation of principal components as the directions that preserve most of the "volume" of the dataset. Our main contribution is a theorem that shows that the corresponding non-convex program has no spurious local optima, and is therefore amenable to many convex solvers. We also confirm our findings numerically.
Principal Component Analysis (PCA) finds the best linear representation of data, and is an indispensable tool in many learning and inference tasks. Classically, principal components of a dataset are interpreted as the directions that preserve most of its "energy", an interpretation that is theoretically underpinned by the celebrated Eckart-Young-Mirsky Theorem. This paper introduces many other ways of performing PCA, with various geometric interpretations, and proves that the corresponding family of non-convex programs have no spurious local optima, while possessing only strict saddle points. These programs therefore loosely behave like convex problems and can be efficiently solved to global optimality, for example, with certain variants of the stochastic gradient descent. Beyond providing new geometric interpretations and enhancing our theoretical understanding of PCA, our findings might pave the way for entirely new approaches to structured dimensionality reduction, such as sparse PCA and nonnegative matrix factorisation. More specifically, we study an unconstrained formulation of PCA using determinant optimisation that might provide an elegant alternative to the deflating scheme commonly used in sparse PCA.
Low-rank approximations of data matrices are an important dimensionality reduction tool in machine learning and regression analysis. We consider the case of categorical variables, where it can be formulated as the problem of finding low-rank approximations to Boolean matrices. In this paper we give what is to the best of our knowledge the first integer programming formulation that relies on only polynomially many variables and constraints, we discuss how to solve it computationally and report numerical tests on synthetic and real-world data.
Consider finite sequences $X_{[1,n]}=X_1\dots X_n$ and $Y_{[1,n]}=Y_1\dots Y_n$ of length $n$, consisting of i.i.d.\ samples of random letters from a finite alphabet, and let $S$ and $T$ be chosen i.i.d.\ randomly from the unit ball in the space of symmetric scoring functions over this alphabet augmented by a gap symbol. We prove a probabilistic upper bound of linear order in $n^{0.75}$ for the deviation of the score relative to $T$ of optimal alignments with gaps of $X_{[1,n]}$ and $Y_{[1,n]}$ relative to $S$. It remains an open problem to prove a lower bound. Our result contributes to the understanding of the microstructure of optimal alignments relative to one given scoring function, extending a theory begun by the first two authors.
We extend Relative Robust Portfolio Optimization models to allow portfolios to optimize their performance when considered relative to a set of benchmarks. We do this in a minimum volatility setting, where we model regret directly as the maximum difference between our volatility and that of a given benchmark. Portfolio managers are also given the option of computing regret as a proportion of the benchmark’s performance, which is more in line with market practice than other approaches suggested in the literature. Furthermore, we propose using regret as an extra constraint rather than as a brand new objective function, so practitioners can maintain their current framework. We also look into how such a triple optimization problem can be solved or at least approximated for a general class of objective functions and uncertainty and benchmark sets. Finally, we illustrate the benefits of this approach by examining its performance against other common methods in the literature in several equity markets.
In the seminal paper on optimal execution of portfolio transactions, Almgren and Chriss (2001) define the optimal trading strategy to liquidate a fixed volume of a single security under price uncertainty. Yet there exist situations, such as in the power market, in which the volume to be traded can only be estimated and becomes more accurate when approaching a specified delivery time. In this paper, we develop a model that accounts for volume uncertainty and we show that a risk-averse trader has benefit in delaying their trades. More precisely, we argue that the optimal strategy is a trade-off between early and late trades in order to balance risk associated with both price and volume. By incorporating a risk term related to the volume to trade, the static optimal strategies suggested by our model avoid the explosion in the algorithmic complexity usually associated with dynamic programming solutions, all the while yielding competitive performance.
Sven Leyffer合作论文数Mathematics and Computer Science Division at Argonne National Laboratory1