
In this paper, we study a queueing process for which the dynamics are changed once the workload in the queue exceeds a predefined threshold, and these new dynamics stay in force until the queue is emptied, at which point the previous dynamics are again reinstalled. For general spectrally positive Lévy processes in each case, we derive expressions for the workload at an independent exponentially distributed random time horizon and study in detail properties of the workload in stationarity. We work out explicit formulas as well as expressions for the optimal changing threshold for special cases of the underlying process dynamics and a chosen set of involved switching, holding, and service cost functions. Funding: The research of O. Boxma and M. Mandjes has been partly funded by the Nederlandse Organisatie voor Wetenschappelijk Onderzoek Gravitation [Project Networks, Grant 024.002.003]. The research of O. Kella is partially funded by Israel Science Foundation [Grant 3336/24] and the Vigevani Chair in Statistics.
In this paper, we study the fluid limit of many-server queues with abandonment (with traffic intensity [Formula: see text]) via a class of nonlinear Volterra equations. For a broad class of service time distributions, we establish the asymptotic behavior of the solutions to the class of nonlinear Volterra equations, which in turn implies the large time behavior of the fluid limit of the many-server queues with abandonment.
This paper studies the applications of state-dependent importance sampling in pricing exotic rainbow options. We demonstrate that for many rainbow options, efficient dynamic importance sampling schemes can be constructed from subsolutions to appropriate partial differential equations. We also introduce an alternative large deviations scaling, which leads to universal importance sampling for rainbow options.
We study optimal coordination in quality-driven rework systems, where each station contributes to an item’s quality and the likelihood of rework depends on these contributions through a production function. Our model links quality targets to processing times via Brownian hitting times and aggregates quality into rework probability. By setting quality targets, a planner jointly shapes processing times and rework rates, determining overall workload. We characterize the attainable workload region and use it to minimize effort costs in two-station networks. The optimal policy depends on (i) complementarity versus substitutability of quality contributions, (ii) each station’s quality impact, and (iii) relative operational costs. We identify three operating regimes: two corner regimes (one active station) and a network regime (both active). As complementarity increases, effort shifts toward the more cost-effective station, whereas finite capacity modifies this threshold, pushing effort toward the less constrained station. The analysis highlights how complementarity shapes efficiency trade-offs and capacity requirements. To illustrate the model’s practical relevance, we calibrate it using published operational statistics from industrial maintenance settings, in which equipment undergoes repair followed by preventive service while offline. The results show that the optimal design is robust to parameter misspecification, achieves substantial cost savings relative to benchmark policies, and can be estimated from standard operational data. Funding: Partial financial support was provided by the Israel Science Foundation [Grant 277/21] and the Bernard M. Gordon Center for Systems Engineering at the Technion. Supplemental Material: The online appendix is available at https://doi.org/10.1287/stsy.2025.0111 .
We analyze the efficiency of parallelization and restart mechanisms for stochastic simulations in model-free settings, where the underlying system dynamics are unknown. Such settings are common in Reinforcement Learning (RL) and rare event estimation, where standard variance-reduction techniques like importance sampling are inapplicable. Focusing on the challenge of reaching rare states under a finite computational budget, we model exploration via random walks and Lévy processes. Based on rigorous probability analysis, our work reveals a phase transition in the success probability as a function of the number of parallel simulations: an optimal number N^* exists, balancing exploration diversity and time allocation per simulation. Beyond this threshold, performance degrades exponentially. Furthermore, we demonstrate that a restart strategy, which reallocates resources from stagnant trajectories to promising regions, can yield an exponential improvement in success probability. In the context of RL, these strategies can improve policy gradient methods by enabling more efficient state-space exploration, leading to more accurate policy gradient estimates.
We show that the state spaces of multifactor Markovian processes, coming from approximations of nonnegative Volterra processes, are given by explicit linear transformation of the nonnegative orthant. We demonstrate the usefulness of this result for applications, including simulation schemes and PDE methods for nonnegative Volterra processes.
Motivated by applications in service systems, we consider queueing systems where each customer must be handled by a server with the right skill set. We focus on optimizing the routing of customers to servers in order to maximize the total payoff of customer–server matches. In addition, customer–server dependent payoff parameters are assumed to be unknown a priori. We construct a machine learning algorithm that adaptively learns the payoff parameters while maximizing the total payoff and prove that it achieves polylogarithmic regret. Moreover, we show that the algorithm is asymptotically optimal up to logarithmic terms by deriving a regret lower bound. The algorithm leverages the basic feasible solutions of a static linear program as the action space. The regret analysis overcomes the complex interplay between queueing and learning by analyzing the convergence of the queue length process to its stationary behavior. We also demonstrate the performance of the algorithm numerically, and have included an experiment with time-varying parameters highlighting the potential of the algorithm in non-static environments.
In this paper, we analyze the optimal management of local memory systems, using the tools of stationary point processes. We provide a rigorous setting of the problem, building upon recent work, and characterize the optimal causal policy that maximizes the hit probability. We specialize the result for the case of renewal request processes and derive a suitable large scale limit as the catalog size N grows to infinity, when a fixed fraction c of items can be stored. We prove that in the limiting regime, the optimal policy amounts to comparing the stochastic intensity (observed hazard rate) of the process with a fixed threshold, defined by a quantile of an appropriate limit distribution, and derive asymptotic performance metrics, as well as sharp estimates for the pre-limit case. Moreover, we establish a connection with optimal timer based policies for the case of monotonic hazard rates. We also present detailed validation examples of our results, including some close form expressions for the miss probability that are compared to simulations. We also use these examples to exhibit the significant superiority of the optimal policy for the case of regular traffic patterns.
We consider an outward degenerate drifted Brownian motion in the quarter plane with oblique reflections on the boundaries. In this article, we explicitly compute the Laplace transforms of the Green's functions associated with the process. These Laplace transforms are expressed as an infinite sum of products using the compensation method. We also derive the asymptotics of the Green's functions along all possible paths and determine the (minimal) Martin boundary. Finally, we provide explicit formulae for all the corresponding harmonic functions.
Our work is part of the close link between continuous-time dissipative dynamical systems and optimization algorithms, and more precisely here, in the stochastic setting. We aim to study stochastic convex minimization problems through the lens of stochastic inertial differential inclusions that are driven by the subgradient of a convex objective function. This will provide a general mathematical framework for analyzing the convergence properties of stochastic second-order inertial continuous-time dynamics involving vanishing viscous damping and measurable stochastic subgradient selections. Our chief goal in this paper is to develop a systematic and unified way that transfers the properties recently studied for first-order stochastic differential equations to second-order ones involving even subgradients in lieu of gradients. This program will rely on two tenets: time scaling and averaging, following an approach recently developed in the literature by one of the co-authors in the deterministic case. Under a mild integrability assumption involving the diffusion term and the viscous damping, our first main result shows that almost surely, there is weak convergence of the trajectory towards a minimizer of the objective function and fast convergence of the values and gradients. We also provide a comprehensive complexity analysis by establishing several new pointwise and ergodic convergence rates in expectation for the convex, strongly convex, and (local) Polyak-Lojasiewicz case. Finally, using Tikhonov regularization with a properly tuned vanishing parameter, we can obtain almost sure strong convergence of the trajectory towards the minimum norm solution.
In many stochastic service systems, decision-makers find themselves making a sequence of decisions, with the number of decisions being unpredictable. To enhance these decisions, it is crucial to uncover the causal impact these decisions have through careful analysis of observational data from the system. However, these decisions are not made independently, as they are shaped by previous decisions and outcomes. This phenomenon is called sequential bias and violates a key assumption in causal inference that one person's decision does not interfere with the potential outcomes of another. To address this issue, we establish a connection between sequential bias and the subfield of causal inference known as dynamic treatment regimes. We expand these frameworks to account for the random number of decisions by modeling the decision-making process as a marked point process. Consequently, we can define and identify causal effects to quantify sequential bias. Moreover, we propose estimators and explore their properties, including double robustness and semiparametric efficiency. In a case study of 27,831 encounters with a large academic emergency department, we use our approach to demonstrate that the decision to route a patient to an area for low acuity patients has a significant impact on the care of future patients.
We derive uniform all-time concentration bound of the type ‘for all [Formula: see text] for some [Formula: see text]’ for TD(0) with linear function approximation. We work with online TD learning with samples from a single sample path of the underlying Markov chain. This makes our analysis significantly different from offline TD learning or TD learning with access to independent samples from the stationary distribution of the Markov chain. We treat TD(0) as a contractive stochastic approximation algorithm with both martingale and Markov noises. Markov noise is handled using the Poisson equation, and the lack of almost-sure guarantees on boundedness of iterates is handled using the concept of relaxed concentration inequalities. Funding: The work of V. S. Borkar was supported in part by the S. S. Bhatnagar Fellowship from the Government of India.
Consider a Markovian tandem line with finite intermediate buffers and an equal number of stations and servers. Servers are flexible but noncollaborative, so that a job can be processed by at most one server at any time. When a job is being processed, it can be damaged and wasted depending on the proficiency of the server. We identify the dynamic server assignment policy that maximizes the long-run average throughput of the system with two stations and two servers. We find that the optimal policy is either a single or a double threshold policy on the number of jobs in the buffer, where the thresholds depend on the service rates and defect probabilities of the two servers at the two stations. For larger systems, we show that the optimal policy may involve server idling and that improving the service rate at any station is always beneficial. Finally, we propose heuristic server assignment policies motivated by experimentation for small systems with finite buffers and analysis of larger systems with infinite buffers. Numerical results suggest that our heuristics yield near-optimal performance. Funding: This research was supported by the National Science Foundation [Grants CMMI-1536990 and CMMI-2127778]. S. Andradóttir was also supported by the National Science Foundation [Grant CMMI-2348409].
We consider the load-balancing system under Poisson arrivals, exponential services, and homogeneous servers under the power-of-d choices routing algorithm, which chooses the queue with minimum length among d randomly sampled queues. We study this system in the sub-Halfin-Whitt regime. In particular, we consider a sequence of systems with n servers, where the arrival rate of the nth system is [Formula: see text] for [Formula: see text]. It was shown that under power-of-d choices routing with [Formula: see text], the queue length behaves similarly to that of joining the shortest queue and that there are asymptotically zero queueing delays. The focus of this paper is to characterize the behavior when d is below this threshold. We obtain high probability bounds on the queue lengths for various values of d and large enough n. In particular, we show that when d is [Formula: see text] for some integer [Formula: see text], then the asymptotic queue length is m with high probability. Moreover, if d grows poly-log in n, that is, slower than any polynomial, but is at least [Formula: see text], the queue length blows up to infinity asymptotically. We obtain these results by using an iterative state space collapse approach. Funding: This work was partially supported by the National Science Foundation [Grants EPCN-2144316 and CMMI-2140534].
We consider a two-dimensional reflected Ornstein–Uhlenbeck (ROU) process that arises as the diffusion approximation for a parallel server network with a randomly split Hawkes arrival process (or a multivariate Hawkes arrival process) in heavy traffic. We study the ergodic properties of the process, including the positive recurrence and rate of convergence in total variation distance and in Wasserstein distance. We also provide a numerical scheme based on a Monte Carlo method to approximate the invariant measure of the process. Funding: G. Pang is partly supported by the National Science Foundation [Grants DMS 2216765 and CMMI 2452829].
Queueing systems present many opportunities for applying machine learning predictions, such as estimated service times, to improve system performance. This integration raises numerous open questions about how predictions can be effectively leveraged to improve scheduling decisions. Recent studies explore queues with predicted service times, typically aiming to minimize job time in the system. We review these works, highlight the effectiveness of predictions, and present open questions on queue performance. We then move to consider an important practical example of using predictions in scheduling, namely large language model (LLM) systems, which presents novel scheduling challenges and highlights the potential for predictions to improve performance. In particular, we consider LLMs performing inference. Inference requests (jobs) in LLM systems are inherently complex; they have variable inference times, dynamic memory footprints that are constrained by key-value store memory limitations, and multiple possible preemption approaches that affect performance differently. We provide background on the important aspects of scheduling in LLM systems and introduce new models and open problems that arise from them. We argue that there are significant opportunities for applying insights and analysis from queueing theory to scheduling in LLM systems. Funding: M. Mitzenmacher was supported in part by the National Science Foundation [Grant CCF-2101140]. R. Shahout was supported in part by the Schmidt Futures Initiative and the Zuckerman Institute. M. Mitzenmacher and R. Shahout were supported in part the National Science Foundation [Grants CNS-2107078 and DMS-2023528].
We consider a dynamic system with multiple types of customers and servers. Each type of waiting customer or server joins a separate queue, forming a bipartite graph with customer-side queues and server-side queues. The platform can match the servers and customers if their types are compatible. The matched pairs then leave the system. The platform will charge a customer a price according to their type when they arrive and will pay a server a price according to their type. The arrival rate of each queue is determined by the price according to some unknown demand or supply functions. Our goal is to design pricing and matching algorithms to maximize the profit of the platform with unknown demand and supply functions, while keeping queue lengths of both customers and servers below a predetermined threshold. This system can be used to model two-sided markets such as ride-sharing markets with passengers and drivers. The difficulties of the problem include simultaneous learning and decision making, and the tradeoff between maximizing profit and minimizing queue length. We use a longest-queue-first matching algorithm and propose a learning-based pricing algorithm, which combines gradient-free stochastic projected gradient ascent with bisection search. We prove that our proposed algorithm yields a sublinear regret Õ(T^5/6) and queue-length bound Õ(T^2/3), where T is the time horizon. We further establish a tradeoff between the regret bound and the queue-length bound: Õ(T^1-γ/4) versus Õ(T^γ) for γ∈ (0, 2/3].
We consider a system of $N$ particles whose interactions are characterized by a (weighted) graph $G^N$. Each particle is a node of the graph with an internal state. The state changes according to Markovian dynamics that depend on the states and connection to other particles. We study the limiting properties, focusing on the dense graph regime, where the number of neighbors of a given node grows with $N$. We show that when $G^N$ converges to a graphon $G$, the behavior of the system converges to a deterministic limit, the graphon mean field approximation. We obtain convergence rates depending on the system size $N$ and cut-norm distance between $G^N$ and $G$. We apply the results for two subcases: When $G^N$ is a discretization of the graph $G$ with individually weighted edges; when $G^N$ is a random graph obtained through edge sampling from the graphon $G$. In the case of weighted interactions, we obtain a bound of order $O(1/N)$. In the random graph case, the error is of order $O(\sqrt{\log(N)/N})$ with high probability. We illustrate the applicability of our results and the numerical efficiency of the approximation through two examples: a graph-based load-balancing model and a heterogeneous bike-sharing system.
The join-the-shortest-queue (JSQ) load-balancing scheme is known to minimize the average response time of jobs in homogeneous systems with identical servers. However, for heterogeneous systems with servers having different processing speeds, finding an optimal load balancing scheme remains an open problem for finite system sizes. Recently, for systems with heterogeneous servers, a variant of JSQ scheme, called the speed-aware-join-the-shortest-queue (SA-JSQ) scheme, has been shown to achieve asymptotic optimality in the fluid-scaling regime where the number of servers n tends to infinity but the normalized the arrival rate of jobs remains constant. In this paper, we show that the SA-JSQ scheme is also asymptotically optimal for heterogeneous systems in the Halfin-Whitt traffic regime where the normalized arrival rate scales are [Formula: see text]. Our analysis begins by establishing that an appropriately scaled and centered version of the Markov process describing system dynamics weakly converges to a two-dimensional reflected Ornstein-Uhlenbeck (OU) process. We then show using Stein’s method that the stationary distribution of the underlying Markov process converges to that of the OU process as the system size increases by establishing the validity of interchange of limits. Finally, through coupling with a suitably constructed system, we show that SA-JSQ asymptotically minimizes the diffusion-scaled total number of jobs and the diffusion-scaled number of waiting jobs in the steady state in the Halfin-Whitt regime among all policies that dispatch jobs based on queue lengths and server speeds.
Blockchain and other decentralized databases, known as distributed ledgers, are designed to store information online where all trusted network members can update the data with transparency. The dynamics of ledger's development can be mathematically represented by a directed acyclic graph (DAG). One essential property of a properly functioning shared ledger is that all network members holding a copy of the ledger agree on a sequence of information added to the ledger, which is referred to as consensus and is known to be related to a structural property of DAG called one-endedness. In this paper, we consider a model of distributed ledger with sequential stochastic arrivals that mimic attachment rules from the IOTA cryptocurrency. We first prove that the number of leaves in the random DAG is bounded by a constant infinitely often through the identification of a suitable martingale, and then prove that a sequence of specific events happens infinitely often. Combining those results we establish that, as time goes to infinity, the IOTA DAG is almost surely one-ended.