We study the variance of Q-function estimators in finite-horizon, finite-state Markov decision process (MDP) tree search. We show that the variance decomposes into three components attributed to the immediate reward collected, probabilistic state transitions, and uncertainty in future state value function estimates. Using this decomposition, we show that the sample variance estimator based on the assumption of i.i.d. paths is biased, underestimating the true variance, and the bias does not vanish in the limit. We then propose a recursive variance estimator that is consistent. To enable efficient storage and computation, we derive an equivalent implementation of the recursive estimator using only node-local statistics that can be iteratively updated. This consistent variance estimator is integrated into two Monte Carlo Tree Search (MCTS) sampling procedures for finite-horizon MDPs. In numerical examples from inventory control and kidney paired donation matching, the new estimator improves the performance of the MCTS algorithm relative to a baseline that uses the i.i.d.-based sample variance estimator.
To address sensitivity analysis and optimization for a discrete-time stochastic epidemic model, we derive unbiased gradient estimators that accommodate uncertainties represented as distributions over the parameters of interest, such as those arising from Bayesian calibration. Specifically, we estimate the sensitivity of total infections over a finite time horizon with respect to the proportion immunized (v) and the contact rate (β). Comparing the proposed estimators with deterministic limit approximations based on large populations reveals differences due to the finite population and time horizon. The estimators exhibit lower variance than finite-difference estimators for the derivative with respect to β, but higher variance for the derivative with respect to v. Simulation experiments indicate parameter uncertainty reduces sensitivity to the parameters of interest. In particular, indirect effects of vaccination, such as herd immunity, are less pronounced compared to when parameters are known. For optimization problems balancing intervention and infection costs, incorporating parametric uncertainty leads to more conservative policies.
We study the problem of generating graphs with prescribed degree sequences for bipartite, directed, and undirected networks. We first propose a sequential method for bipartite graph generation and establish a necessary and sufficient interval condition that characterizes the admissible number of connections at each step, thereby guaranteeing global feasibility. Based on this result, we develop bipartite graph enumeration and sampling algorithms suitable for different problem sizes. We then extend these bipartite graph algorithms to the directed and undirected cases by incorporating additional connection constraints, as well as feasibility verification and symmetric connection steps, while preserving the same algorithmic principles. Finally, numerical experiments demonstrate the performance of the proposed algorithms, particularly their scalability to large instances where existing methods become computationally prohibitive.
Leibniz derivative estimation is a Monte Carlo technique for estimating derivatives of a discontinuous sample performance in stochastic models with respect to parameters of interest. By combining the push-out likelihood ratio (LR) method with Leibniz integral rules, it generalizes a broad class of existing LR-based derivative estimators. However, as an LR-based method, its variance is often higher than that of perturbation analysis-based methods and may grow linearly with the dimension of the stochastic input whose distribution depends on the parameter. In this paper, we propose a recursive conditioning approach and combine it with the Leibniz derivative estimation framework. The resulting conditional Leibniz estimator does not involve LR terms and therefore is not subject to variance growth with the input dimension. It also has a simple form and is easy to implement. We apply the method to an American call min-option model, and simulation results show its effectiveness and low-variance performance.
With the relentless increase in computing power and the ubiquitous availability of data in many industries, the fields of simulation optimization and artificial intelligence have emerged at the scientific and engineering forefront in their societal impact, manifested in the pervasiveness of technologies such as large language models, chatbots, digital twins, and agent-based systems. We examine cross-fertilization between simulation optimization and artificial intelligence, with a particular focus on reinforcement learning, highlighting research that has been mutually beneficial. After reviewing relevant methodologies from stochastic simulation and optimization, we provide an overview of foundational approaches of machine learning, and then present examples of the synergies between the fields, followed by real-world applications and case studies, including some futurist and forward-looking concepts.
We establish structural properties of optimal stopping problems under time-consistent dynamic (coherent) risk measures, focusing on value function monotonicity and the existence of control limit (threshold) optimal policies. While such results are well developed for risk-neutral (expected-value) models, they remain underexplored in risk-averse settings. Coherent risk measures typically lack the tower property and are subadditive rather than additive, complicating structural analysis. We show that value function monotonicity mirrors the risk-neutral case. Moreover, if the risk envelope associated with each coherent risk measure admits a minimal element, the risk-averse optimal stopping problem reduces to an equivalent risk-neutral formulation. We also develop a general procedure for identifying control limit optimal policies and use it to derive practical, verifiable conditions on the risk measures and MDP structure that guarantee their existence. We illustrate the theory and verify these conditions through optimal stopping problems arising in operations, marketing, and finance.
This tutorial addresses the use of stochastic gradients in simulation optimization, including methodology and algorithms, theoretical convergence analysis, and applications. Specific topics include stochastic gradient estimation techniques - both direct unbiased and indirect finite-difference-based; stochastic approximation algorithms and their convergence rates; stochastic gradient descent (SGD) algorithms with momentum and reusing past samples via importance sampling; and a real-world application in geographical partitioning.
We review simulation optimization methods and their connection with artificial intelligence (AI) techniques. In particular, we focus on two areas: stochastic gradient estimation, which plays a central role in training neural networks for deep learning and reinforcement learning; and ranking and selection, which can be used as the node selection policy in Monte Carlo tree search. We also review the literature on inventory management, which has been studied in both simulation optimization and AI.
Importance sampling (IS) is a technique that enables statistical estimation of output performance at multiple input distributions from a single nominal input distribution. IS is commonly used in Monte Carlo simulation for variance reduction and in machine learning applications for reusing historical data, but its effectiveness can be challenging to quantify. In this work, we establish a new result showing the tightness of polynomial concentration bounds for classical IS likelihood ratio (LR) estimators in certain settings. Then, to address a practical statistical challenge that IS faces regarding potentially high variance, we propose new truncation boundaries when using a truncated LR estimator, for which we establish upper concentration bounds that imply an exponential convergence rate. Simulation experiments illustrate the contrasting convergence rates of the various LR estimators and the effectiveness of the newly proposed truncation-boundary LR estimators for examples from finance and machine learning.
We develop a novel stochastic derivative estimation framework for sample performance functions that are discontinuous in the parameter of interest, based on the multidimensional Leibniz integral rule. When discontinuities arise from indicator functions, we embed the indicator functions into the sample space, yielding a continuous performance function over a parameter-dependent domain. Applying the Leibniz integral rule in this case produces a single-run, unbiased derivative estimator. For general discontinuous functions, we apply a change of variables to shift parameter dependence into the sample space and the underlying probability measure. Applying the Leibniz integral rule leads to two terms: a standard likelihood ratio (LR) term from differentiating the underlying probability measure and a surface integral from differentiating the boundary of the domain. Evaluating the surface integral may require simulating multiple sample paths. Our proposed Leibniz integration framework generalizes the generalized LR (GLR) method and provides intuition as to when the surface integral vanishes, thereby enabling single-run, easily implementable estimators. Numerical experiments demonstrate the effectiveness and robustness of our methods.
This tutorial serves as an introductory guide to Monte Carlo tree search (MCTS), a versatile methodology for sequential decision making under uncertainty through stochastic/Monte Carlo simulation. MCTS gained notoriety from its pivotal role in Google DeepMind's AlphaZero and AlphaGo, hailed as major breakthroughs in artificial intelligence (AI) due to AlphaGo defeating the reigning human world Go champion Lee Sedol in 2016 and the world's top-ranked Go player Ke Jie in 2017. AlphaZero, without requiring any domain-specific knowledge beyond the game rules (tabula rasa), achieved remarkable success by surpassing previous benchmarks in Go and outperforming leading AI opponents in chess (Stockfish) and shogi (Elmo) after just 24 hours of MCTS-driven reinforcement learning. We demonstrate the building blocks of MCTS and its performance through decision trees and the game of Othello, and provide an empirical simulation study for the latter.
Users of escalators and moving walkways with sufficient space to accommodate two lanes of users often follow an implied etiquette for the two lanes: one for walking and one for standing (in the US, China, and many countries, the convention is ``walk left, stand right"). When there is high volume, e.g., when exiting a subway or train station or at the conclusion of an athletic event, the escalators often experience bottleneck congestion that constitutes the primary source of delay. It has been suggested that during such high-congestion periods it would be more efficient if everyone just stood in both lanes, with empirical evidence used to support this counterintuitive finding. Simple deterministic queueing models are used to show under what conditions such results hold and also to provide further insights regarding tradeoffs between performance metrics as a function of the distribution of walkers and standers, which could inform practical implementation policies to increase efficiency.
Finite-Time Performance of Gradient-Based Stochastic Approximation Algorithms Stochastic approximation (SA) algorithms are among the most commonly used methods for addressing stochastic optimization problems. In “Technical Note: On the Convergence Rate of Stochastic Approximation for Gradient-Based Stochastic Optimization,” J. Hu and M. C. Fu provide a new error bound for SA algorithms when applied to a class of problems with convex differentiable structures. The bound allows a detailed characterization of an algorithm’s finite-time performance through directly analyzing the bias and variance of the employed gradient estimator. The main result is then applied to study and compare the efficiency of traditional Kiefer-Wolfowitz algorithms and those based on random direction finite-difference gradient estimators such as simultaneous perturbation SA. The analysis leads to new insights into when it may be advantageous to use such randomized gradient estimates under various problem and algorithm parameter settings.
AboutSectionsRequest Access ToolsAdd to favoritesDownload CitationsTrack CitationsPermissionsReprints ShareShare onFacebookTwitterLinked InEmail Go to Section HomeINFORMS TutORials in Operations ResearchTutorials in Operations Research: Advancing the Frontiers of OR/MS: From Methodologies to Applications Simulation Optimization in the New Era of AIYijie Peng, Chun-Hung Chen, Michael C. FuYijie Peng, Chun-Hung Chen, Michael C. FuPublished Online:13 Oct 2023https://doi.org/10.1287/educ.2023.0264AbstractWe review simulation optimization methods and discuss how these methods underpin modern artificial intelligence (AI) techniques. In particular, we focus on three areas: stochastic gradient estimation, which plays a central role in training neural networks for deep learning and reinforcement learning; simulation sample allocation, which can be used as the node selection policy in Monte Carlo tree search; and variance reduction, which can accelerate training procedures in AI.Funding: This work was supported in part by the U.S. Air Force Office of Scientific Research [Grant FA95502010211], the National Natural Science Foundation of China [Grants 72250065, 72022001, and 71901003], and U.S. National Science Foundation [Grant FAIN 2123683]. Your Access Options Login Options INFORMS Member Login Nonmember Login Purchase Options Save for later Item saved, go to cart Tutorials in OR, TutorialsNew $20.00 Add to cart Tutorials in OR, TutorialsNew Checkout Other Options Token Access Insert token number Claim access using a token Restore guest access Applies for purchases made as a guest Previous Back to Top Next FiguresReferencesRelatedInformation Tutorials in Operations Research: Advancing the Frontiers of OR/MS: From Methodologies to ApplicationsOctober 2023 Article Information Metrics Information Published Online:October 13, 2023 Copyright © 2023, INFORMSCite asYijie Peng, Chun-Hung Chen, Michael C. Fu (2023) Simulation Optimization in the New Era of AI. INFORMS TutORials in Operations Research null(null):82-108. https://doi.org/10.1287/educ.2023.0264 Keywordssimulation optimizationgradient estimationsimulation sample allocationvariance reductiondeep learningreinforcement learningMonte Carlo tree searchPDF download
Recently, automated vulnerability repair (AVR) approaches have been widely adopted to combat the increasing number of software security issues. In particular, transformer-based models achieve competitive results. While existing models are learned to generate vulnerability repairs, existing AVR models lack a mechanism to provide their models with the precise location of vulnerable code (i.e., models may generate repairs for the non-vulnerable areas). To address this problem, we base our framework on the VIT-based approaches for object detection that learn to locate bounding boxes via the cross-matching between object queries and image patches. We cross-match vulnerability queries and their corresponding vulnerable code areas through the cross-attention mechanism to generate more accurate repairs. To strengthen our cross-matching, we propose to learn a novel vulnerability query mask that greatly focuses on vulnerable code areas and integrate it into the cross-attention. Moreover, we also incorporate the vulnerability query mask into the self-attention to learn embeddings that emphasize the vulnerable areas of a program. Through an extensive evaluation using the real-world 5,417 vulnerabilities, our approach outperforms all of the baseline methods by 3.39%-33.21%. The training code and pre-trained models are available at https://github.com/AVR-VQR/VQR.
We consider a stopping problem and its application to the decision-making process regarding the optimal timing of organ transplantation for individual patients. At each decision period, the patient state is inspected and a decision is made whether to transplant. If the organ is transplanted, the process terminates; otherwise, the process continues until a transplant happens or the patient dies. Under suitable conditions, we show that there exists a control limit optimal policy. We propose a smoothed perturbation analysis (SPA) estimator for the gradient of the total expected discounted reward with respect to the control limit. Moreover, we show that the SPA estimator is asymptotically unbiased.
We review the literature on individual patient organ acceptance decision making by presenting a Markov Decision Process (MDP) model to formulate the organ acceptance decision process as a stochastic control problem. Under the umbrella of the MDP framework, we classify and summarize the major research streams and contributions. In particular, we focus on control limit-type policies, which are shown to be optimal under certain conditions and easy to implement in practice. Finally, we briefly discuss open problems and directions for future research.
The objective in a traditional reinforcement learning (RL) problem is to find a policy that optimizes the expected value of a performance metric such as the infinite-horizon cumulative discounted or long-run average cost/reward. In practice, optimizing the expected value alone may not be satisfactory, in that it may be desirable to incorporate the notion of risk into the optimization problem formulation, either in the objective or as a constraint. Various risk measures have been proposed in the literature, e.g., exponential utility, variance, percentile performance, chance constraints, value at risk (quantile), conditional value-at-risk, prospect theory and its later enhancement, cumulative prospect theory. In this book, we consider risk-sensitive RL in two settings: one where the goal is to find a policy that optimizes the usual expected value objective while ensuring that a risk constraint is satisfied, and the other where the risk measure is the objective. We survey some of the recent work in this area specifically where policy gradient search is the solution approach. In the first risk-sensitive RL setting, we cover popular risk measures based on variance, conditional value-at-risk, and chance constraints, and present a template for policy gradient-based risk-sensitive RL algorithms using a Lagrangian formulation. For the setting where risk is incorporated directly into the objective function, we consider an exponential utility formulation, cumulative prospect theory, and coherent risk measures. This non-exhaustive survey aims to give a flavor of the challenges involved in solving risk-sensitive RL problems using policy gradient methods, as well as outlining some potential future research directions.
We analyze a tree search problem with an underlying Markov decision process, in which the goal is to identify the best action at the root that achieves the highest cumulative reward. We present a new tree policy that optimally allocates a limited computing budget to maximize a lower bound on the probability of correctly selecting the best action at each node. Compared to widely used upper confidence bound (UCB) tree policies, the new tree policy presents a more balanced approach to manage the exploration and exploitation tradeoff when the sampling budget is limited. Furthermore, UCB assumes that the support of reward distribution is known, whereas our algorithm relaxes this assumption. Numerical experiments demonstrate the efficiency of our algorithm in selecting the best action at the root.
The generalized likelihood ratio (GLR) method is a recently introduced gradient estimation method for handling discontinuities in a wide range of sample performances. We put the GLR methods from previous work into a single framework, simplify regularity conditions to justify the unbiasedness of GLR, and relax some of those conditions that are difficult to verify in practice. Moreover, we combine GLR with conditional Monte Carlo methods and randomized quasi-Monte Carlo methods to reduce the variance. Numerical experiments show that variance reduction could be significant in various applications.