
This paper is a study on solutions of the sample average approximation method to solve compound stochastic programs. We derive nonasymptotic upper estimates for probabilities of the approximation errors. The results depend on the sample size and use explicit terms rather than unspecified universal constants. They allow for immediate conclusions about nonasymptotic rates for the optimal solutions. Additionally, the results can be used to construct nonasymptotic confidence regions for solutions of compound stochastic programs. In the special case of classical risk-neutral stochastic programs, we end up with upper estimates of deviation probabilities for M-estimators, and their nonasymptotic rates. Moreover, we may also demonstrate how to apply the results to sample average approximation of risk-averse stochastic programs. In this respect we consider stochastic programs expressed in terms of absolute semideviation risk measures and Average Value at Risk. The investigations are based on concentration inequalities from a recent contribution by the author. The line of reasoning does not rely on pathwise analytical properties of the objectives. In particular, continuity or convexity in the parameter is not imposed in advance as usual in the literature on the sample average approximation method. The main results also apply to objectives with Hölder continuous paths. Moreover, they also work for objectives whose paths are piecewise Hölder continuous, as, for example, in two-stage mixed-integer programs.
We investigate a pricing rule that is applicable for streams of income or contingent claim liabilities and study how this rule changes under additional insider-type information that an investor might obtain. Considering a model where the risky asset might have jumps, we obtain an explicit form of the associated state price density for the three different types of agents considered in Ernst and Rogers: one who has no information about the jumps, one who knows in advance exactly when each jump will occur, and one who has no information about the time of the jumps but has partial information about the size of each jump. For each of these agents, we provide characterizations of the pricing rule and establish a representation formula, allowing us to quantify the value of partial information for streams of labor income or contingent claim liabilities. Our work is motivated by finding and characterizing a pricing rule that, both with or without partial information about jumps, assigns different values of information for different income streams or contingent claim liabilities.
The principal contribution of this paper consists in developing a wide range of algorithmic ideas and analytical insights around the continuous-time joint replenishment problem, culminating in a deterministic framework for efficiently approximating optimal dynamic policies to any desired level of accuracy. These advances enable us to derive a compactly encoded replenishment policy whose long-run average cost is within factor 1 + epsilon of the dynamic optimum, arriving at an efficient polynomial-time approximation scheme (EPTAS). Technically speaking, our approach hinges on affirmative resolutions to two fundamental open questions. In relation to feasibility of scalable discretization, we devise the first efficient discretization-based framework for approximating the joint replenishment problem. Specifically, we prove that every continuous-time infinite-horizon instance can be reduced to a corresponding discrete-time O(n3 )-period instance, while incurring epsilon 6 a multiplicative optimality loss of at most 1 + epsilon. Then, in regard to enhanced guarantees for the discrete setting, we improve on the O(22O(1=epsilon) & centerdot; (nT)O(1))-time approximation scheme of Nonner and Sviridenko (IPCO '13) for the discrete-time joint replenishment problem. Beyond an exponential improvement in running time, we demonstrate that the key pillars of their methodology-randomization and hierarchical decompositions-can be entirely avoided, while concurrently offering a streamlined analysis.
Sion and Wolfe [Sion M, Wolfe P (1957) On a game without a value. Contributions to the Theory of Games III, Annals of Mathematics Studies (Princeton University Press, Princeton, NJ), 299-306.] presented a two-person zero-sum game on the unit square without a value. In the present paper, we analyze finite-grid approximations of the Sion-Wolfe game. We find that, as the number of grid points tends to infinity and the payoff function approaches that of the infinite game, the limiting value of finite approximations may lie within, on the boundary of, or even outside the interval defined by the lower and upper values of the infinite game. Although these discrepancies can be explained, our findings underscore the need for great care, even in the case of two-person zero-sum games, when using finite approximations for the analysis of infinite games. Open Access Statement: This work is licensed under a Creative Commons Attribution 4.0 International License. You are free to copy, distribute, transmit and adapt this work, but you must attribute this work as "Mathematics of Operations Research. Copyright (c) 2026 The Author(s). https://doi.org/10. 1287/moor.2022.0254, used under a Creative Commons Attribution License: https://creativecommons. org/licenses/by/4.0/."
We study a class of stochastic exchangeable teams with a finite number of decision makers (DMs) as well as their mean-field limits with infinitely many DMs. In the finite-population regime, we study exchangeable teams under the centralized information structure. For the infinite-population setting, we study both the centralized information structure and the decentralized mean-field information-sharing structure. The paper makes the following main contributions. (i) For finite-population exchangeable teams, we establish the existence of an optimal policy that is exchangeable (permutation invariant) and Markovian. (ii) As our main result in the paper, we show that a sequence of exchangeable optimal policies for finite-population settings (which satisfies a measure-valued Markov decision problem (MDP) formulation (following a work by Ba & uml;uerle) converges to a decentralized symmetric (identical) and conditionally independent (given the mean-field) policy for the infinite-population problem, which is then globally optimal under both the centralized information structure as well as the mean-field-sharing information structure. (iii) This result establishes the existence of a symmetric, independent, decentralized optimal randomized policy for the infinite-population problem and proves the optimality of the limiting measure-valued MDP for the representative DM. Our paper thus establishes the relation between controlled McKean-Vlasov dynamics and the optimal infinite-population decentralized stochastic control problem (without an a priori restriction of symmetry in policies of individual agents) for the first time to our knowledge (beyond several special cases). We also establish near optimality of a numerical method for solving this problem. (iv) Finally, we show that symmetric, independent, decentralized optimal randomized policies are approximately optimal for the corresponding finite-population team with a large number of DMs under the centralized information structure.
This paper develops a framework that integrates finite-sample statistical learning with queueing asymptotic analysis to design matching policies. The stochastic matching model we consider assumes heterogeneous demand (customers) and heterogeneous supply (workers) arrive randomly over time, each with a randomly sampled patience time, and are lost (renege) if forced to wait longer than that time to be matched. Because the interarrival and patience-time distributions are unknown, matching decisions must be made based on historical (offline) data. We leverage asymptotic analysis to formulate a deterministic, data-driven (fluid) matching problem (DDMP) that approximates the original stochastic matching problem. We establish finite-sample statistical guarantees on the objective value gap between the DDMP solution and the ground-truth matching problem solution, which requires a novel uniform error bound involving the patience-time quantile function. We show that a discrete-review, estimate-then-match-type policy is epsilon-asymptotically optimal with high probability as arrival rates grow large.
This paper considers the stochastic symmetric cone linear complementarity problem (S-SCLCP), which includes the stochastic linear complementarity problem and the stochastic second-order cone linear complementarity problem as special cases. We propose a new expected residual minimization (ERM) formulation for S-SCLCP and apply the Monte Carlo technique to generate the corresponding approximation problem. Different from existing ERM formulations for stochastic complementarity problems, the proposed ERM formulation is based on a smooth C-function and a variant merit function. Relying on the Euclidean Jordan algebra associated with symmetric cones, we address several important issues, including coerciveness, existence of solutions, global convergence, and exponential convergence rate. Furthermore, we present some numerical examples and the practical applications of ERM schemes in solving an uncertain Nash-Cournot game and a stochastic optimal power flow problem in the radial network, demonstrating the effectiveness of this method.
Adaptive submodularity is a fundamental concept in stochastic optimization, with numerous applications such as sensor placement, hypothesis identification, and viral marketing. We consider the problem of covering an adaptive submodular function at minimum expected cost, where the random realizations of different items may be correlated. We show that the natural greedy policy has an approximation ratio of 4 center dot (1 + ln Q), where Q is the goal value. We also show that the greedy policy has approximation ratio of at least 1:3 center dot (1 + ln Q) even when Q = 1, which invalidates a prior result on adaptive submodular cover. Moreover, we consider a significantly more general objective of minimizing the pth moment of the coverage cost and show that the greedy policy simultaneously achieves a (p + 1)(p+1) center dot (ln Q + 1)(p) approximation guarantee for all p >= 1. All our approximation ratios are best possible up to constant factors (assuming P not equal NP). Our results also extend to the setting where one wants to cover multiple adaptive submodular functions, for which we obtain the same approximation guarantees.
This paper investigates an optimal consumption, investment, and early retirement problem in the presence of a mandatory retirement date and a borrowing constraint that prohibits borrowing against future labor income during employment. To handle the borrowing constraint, we employ a dual-martingale approach and reformulate the problem as a finite-horizon, two-player, zero-sum game between a singular controller and a stopper. The value of the game is characterized by a parabolic variational inequality with both obstacle and gradient constraints, giving rise to two time-dependent free boundaries that determine the optimal retirement threshold and the wealth-binding region. Using advanced PDE techniques and nonstandard analytical arguments, we prove the existence and uniqueness of a strong solution to the variational inequality and derive key properties of the free boundaries, including their monotonicity and smoothness. We further establish that the solution coincides with the game value and provide a duality theorem that characterizes the optimal strategy. To the best of our knowledge, this is the first study in the mathematical finance literature to analyze a finite-horizon, zero-sum game involving both singular control and stopping, offering novel insights into the joint dynamics of retirement timing, consumption, and portfolio choice.
We consider linear stochastic approximation (LSA) with constant step size and Markovian data. Viewing the joint process of the data and LSA iterate as a time-homogeneous Markov chain, we prove its convergence to a unique limiting and stationary distribution and establish nonasymptotic, geometric convergence rates. Furthermore, we show that the bias vector of this limit admits an infinite series expansion with respect to the step size. Consequently, the bias is proportional to the step size up to higher-order terms. This result contrasts with LSA under independent and identically distributed data, for which the bias vanishes. In the reversible chain setting, we characterize the relationship between the bias and the mixing time of the Markovian data, establishing that they are roughly proportional to each other. Although Polyak-Ruppert tail averaging reduces the variance of LSA iterates, it does not affect the bias. The above characterization implies that the bias can be reduced using Richardson-Romberg extrapolation with m >= 2 step sizes, eliminating the m(-1) leading terms in the bias expansion. This extrapolation scheme leads to an exponentially smaller bias and an improved mean-squared error both theoretically and empirically. Our results immediately apply to the temporal difference learning algorithm with linear function approximation and stochastic gradient descent applied to quadratic functions.
We study envy-freeness up to any good (EFX) in settings where valuations can be represented via a graph of arbitrary size where vertices correspond to agents and edges to items. An item (edge) has zero marginal value to all agents (vertices) not incident to the edge. Each vertex may have an arbitrary monotone valuation on the set of incident edges. We first consider allocations that correspond to orientations of the edges, where we show that EFX does not always exist, and furthermore, that it is NP-complete to decide whether an EFX orientation exists. Our main result is that EFX allocations exist for this setting. This is one of the few cases where EFX allocations are known to exist for more than three agents.
We study sample-efficient reinforcement learning (RL) under the general framework of interactive decision making, which includes the Markov decision process, partially observable Markov decision process, and predictive state representation (PSR) as special cases. We propose a novel complexity measure, the generalized eluder coefficient (GEC), which characterizes the fundamental trade-off between exploration and exploitation in online interactive decision making in the context of function approximation. We show that RL problems with low GECs form a remarkably rich class, which subsumes low Bellman eluder dimension problems, bilinear class, low witness rank problems, Partially observable bilinear (PO-bilinear) class, and generalized regular PSR, where generalized regular PSR, a new tractable PSR class identified by us, includes nearly all known tractable partially observable RL models. Furthermore, in terms of algorithm design, we propose a generic posterior sampling algorithm, which can be implemented in both model-free and model-based fashions, under both fully observable and partially observable settings. We prove that the proposed generic posterior sampling algorithm is sample efficient by establishing a sublinear regret upper bound in terms of the GEC. In summary, we provide a new and unified understanding of both fully observable and partially observable RL.
We study a dynamic bipartite matching market model where agents arrive and depart randomly through Poisson processes. Our proposed mechanisms are for minimizing unmatched agents by determining whom and when to match. Our main contribution is establishing performance bounds for different local mechanisms with varying timing strategies. We find that the Patient algorithm, which delays matching to increase market thickness, outperforms the Greedy algorithm by an exponential factor. Notably, the Patient algorithm requires the planner to identify departing agents, making it an optimal algorithm. Without this requirement, the Greedy algorithm is nearly optimal. We also examine a one-sided market, such as labor or freight exchange markets, where only one side can make decisions. In this scenario, we show that the Greedy and Patient algorithms have similar performance, suggesting that delaying matching time may not be advantageous. This finding contrasts with the bipartite and nonbipartite cases explored in recent literature.
In this paper, the concept of coderivatives at infinity of set-valued mappings is introduced. Well-posedness properties at infinity of set-valued mappings as well as Mordukhovich's criterion at infinity are established. Optimality conditions at infinity in set-valued optimization are also provided. The obtained results, which give new information even in the classical cases of smooth single-valued mappings, provide complete characterizations of the properties under consideration in the setting at infinity of set-valued mappings.
This paper studies consumption-portfolio optimization problems with habit formation in a regime-switching market. The habit level, which reflects the endogenous impact of past consumption, also involves regime switching and jump diffusion. Because of the presence of general utility functions and path-dependent random parameters, we use the market completion method and introduce some additional jump assets to address the problems. After reducing the problems to solving a stochastic Hamilton-Jacobi-Bellman equation, we derive the optimal control by a joint adoption of envelope theorem and backward stochastic partial differential equation. In general, the optimal portfolio strategy includes the demand of jump assets for hedging against the regime-switching and jumpdiffusion risk. In particular, for power/logarithmic utility, we obtain a closed-form solution in the enlarged complete market. For comparison, we also study the power/logarithmic utility case with many specific conditions in the primal incomplete market by restricting the positions of the jump assets to zero.
Nonparametric choice models offer broad applicability and robustness. However, the exponentially large parameter space leads practitioners to use heuristics for estimation. We introduce an alternative approach to modeling and estimating nonparametric choice models using discrete Fourier analysis. We demonstrate that any choice function can be approximated with a small number of Fourier parameters. Our sample-efficient, active-learning algorithms, without requiring an explicit model description, need at most poly(logn, 1=epsilon) data queries to estimate any choice function up to epsilon accuracy. Computational studies show significant error reduction with Fourier methods compared with common heuristics for nonparametric choice estimation in both simulated and real data.
We study an optimal stochastic control problem in which a firm's cash/surplus process is controlled by dividend payments and capital injections. We consider absolutely continuous dividend policies subject to a level-dependent upper bound on the dividend rate and general capital injection strategies. We construct an optimal solution for which either the optimal capital injections consist of a forced bailout strategy when the cash process reaches zero or no injection of capital is ever made and ruin is eventually reached. This gives rise to two distinct dividend optimization problems for which the solutions are shown to be mean-reverting dividend strategies refracted at optimal thresholds. To prove the existence of the optimal threshold in the forced-injection case, we use the theory of viscosity solutions and characterize the optimal threshold in terms of the derivative of the value function. By a uniqueness result for the solution of the associated HJB equation, we show that the value function corresponds to the performance function of a mean-reverting dividend strategy, which we compute explicitly using results from fluctuation theory, and we characterize the optimal threshold. Finally, we give a complete solution to the general problem and characterize the dichotomy by proving a comparison theorem based on the value functions (or their derivatives) at zero for the two dividend optimization problems.
A sampling-based method is introduced to approximate the Gittins index for a general family of alternative bandit processes. The approximation consists of a truncation of the optimization horizon and support for the immediate rewards, an optimal stopping value approximation, and a stochastic approximation procedure. Finite-time error bounds are given for the three approximations, leading to a procedure to construct a confidence interval for the Gittins index using a finite number of Monte Carlo samples as well as an epsilon-optimal policy for the family of alternative bandit processes. Proofs are given for almost sure convergence and a central limit theorem for the sampling-based Gittins index approximation. In a numerical study, the quality of the approximation is verified for the Bernoulli bandit and the Gaussian bandit with known variance, and the method is shown to significantly outperform Thompson sampling and the Bayesian upper-confidencebound algorithms for a novel random effects multi-armed bandit.