Police patrol units need to split their time between performing preventive patrol and being dispatched to serve emergency incidents. In the existing literature, patrol and dispatch decisions are often studied separately. We consider joint optimization of these two decisions to improve police operations efficiency and reduce response time to emergency calls. Methodology/results: We propose a novel method for jointly optimizing multi-agent patrol and dispatch to learn policies yielding rapid response times. Our method treats each patroller as an independent Q-learner (agent) with a shared deep Q-network that represents the state-action values. The dispatching decisions are chosen using mixed-integer programming and value function approximation from combinatorial action spaces. We demonstrate that this heterogeneous multi-agent reinforcement learning approach is capable of learning joint policies that outperform those optimized for patrol or dispatch alone. Managerial Implications: Policies jointly optimized for patrol and dispatch can lead to more effective service while targeting demonstrably flexible objectives, such as those encouraging efficiency and equity in response.
We consider a two-sided marketplace in which a market operator sells services to clients and buys services from vendors. The market operator determines the prices dynamically for both clients and vendors. Services are transacted in discrete units called jobs, and the jobs are characterized by their types, service deadlines, client prices, and vendor prices. Jobs are submitted by clients to the marketplace, and the market operator then lists the available jobs. Vendors can view the available jobs and choose them based on their preferences. We consider an infinite-horizon long-run average reward Markov decision process (MDP) model of the market operator's dynamic pricing problem. The MDP consists of the arrival processes on both sides of the marketplace as well as the choice behavior of both clients and vendors. Because solving this MDP directly is impractical for large market sizes, we study a discrete-time fluid approximation of the problem. This approximation results in a simple pricing policy in which each job's price trajectory depends on the time remaining until the service deadline but does not depend on the other jobs available in the marketplace. We show that this policy is asymptotically optimal with a loss ratio of order O(1=61) on the long-run average reward, where 61 represents the scale of demand and supply. The performance of the fixed price trajectory policy is compared with other heuristics, including a constant pricing policy and a state-dependent pricing policy. We also extend the model to continuous-time and finite-horizon settings.
Crowdsourced delivery platforms operate as an intermediary between consumers who require delivery tasks and couriers who make these deliveries; both of which are uncertain. The main challenge of a crowdsourced delivery platform is to meet a service level for their customers (e.g., 95% on-time delivery) by serving dynamically arriving delivery tasks with time windows. The two critical courier management decisions for a platform are how to schedule couriers and how to assign delivery tasks to couriers. These two decisions can be centralized (i.e., decided by the platform) or decentralized (i.e., decided by the couriers). Centralizing these decisions produces a more reliable workforce while decentralizing them may come with cost savings to the platform and allows more freedom to couriers in deciding when and where to work. Crowdsourced delivery platforms have begun to utilize both courier types simultaneously (i.e., a hybrid system) with the hope of reaping the advantages of each. In this paper, we address the challenge of capacity planning for a crowdsourced delivery platform that utilizes both centralized (committed) and decentralized (ad-hoc) couriers. We present fluid models for delivery systems using either type of courier, and a hybrid formulation. Our theoretical, numerical, and simulation results establish the superiority of a hybrid system over each pure system in a majority of real-world scenarios.
We study the online saddle point problem, an online learning problem where at each iteration, a pair of actions needs to be chosen without knowledge of the current and future (convex-concave) payoff functions. The objective is to minimize the gap between the cumulative payoffs and the saddle point value of the aggregate payoff function, which we measure using a metric called saddle point regret (SP-Regret). The problem generalizes the online convex optimization framework, but here, we must ensure that both players incur cumulative payoffs close to that of the Nash equilibrium of the sum of the games. We propose an algorithm that achieves SP-Regret proportional to [Formula: see text] in the general case, and [Formula: see text] SP-Regret for the strongly convex-concave case. We also consider the special case where the payoff functions are bilinear and the decision sets are the probability simplex. In this setting, we are able to design algorithms that reduce the bounds on SP-Regret from a linear dependence in the dimension of the problem to a logarithmic one. We also study the problem under bandit feedback and provide an algorithm that achieves sublinear SP-Regret. We then consider an online convex optimization with knapsacks problem motivated by a wide variety of applications, such as dynamic pricing, auctions, and crowdsourcing. We relate this problem to the online saddle point problem and establish [Formula: see text] regret using a primal-dual algorithm.
Estimating the total treatment effect (TTE) of a new feature in social platforms is crucial for understanding its impact on user behavior. However, the presence of network interference, which arises from user interactions, often complicates this estimation process. Experimenters typically face challenges in fully capturing the intricate structure of this interference, leading to less reliable estimates. To address this issue, we propose a novel approach that leverages surrogate networks and the pseudo inverse estimator. Our contributions can be summarized as follows: (1) We introduce the surrogate network framework, which simulates the practical situation where experimenters build an approximation of the true interference network using observable data. (2) We investigate the performance of the pseudo inverse estimator within this framework, revealing a bias-variance trade-off introduced by the surrogate network. We demonstrate a tighter asymptotic variance bound compared to previous studies and propose an enhanced variance estimator outperforming the original estimator. (3) We apply the pseudo inverse estimator to a real experiment involving over 50 million users, demonstrating its effectiveness in detecting network interference when combined with the difference-in-means estimator. Our research aims to bridge the gap between theoretical literature and practical implementation, providing a solution for estimating TTE in the presence of network interference and unknown interference structures.
Motivated by diverse applications in sharing economy and online marketplaces, we consider optimal pricing and matching control in a two-sided queueing system. We assume that heterogeneous customers and servers arrive to the system with price-dependent arrival rates. The compatibility between servers and customers is specified by a bipartite graph. Once a pair of customer and server are matched, they depart from the system instantaneously. The objective is to maximize the long-run average profits of the system while minimizing average waiting time. We first propose a static pricing and max-weight matching policy, which achieves O(√ n) optimality rate when all of the arrival rates are scaled by n. We further show that a dynamic pricing and modified max-weight matching policy achieves an improved O(n1/3) optimality rate. In addition, we propose a constraint generation algorithm that solves value function approximation of the MDP and demonstrate strong numerical performance of this algorithm.
In randomized experiments, the classic stable unit treatment value assumption (SUTVA) states that the outcome for one experimental unit does not depend on the treatment assigned to other units. However, the SUTVA assumption is often violated in applications such as online marketplaces and social networks where units interfere with each other. We consider the estimation of the average treatment effect in a network interference model using a mixed randomization design that combines two commonly used experimental methods: Bernoulli randomized design, where treatment is independently assigned for each individual unit, and cluster-based design, where treatment is assigned at an aggregate level. Essentially, a mixed randomization experiment runs these two designs simultaneously, allowing it to better measure the effect of network interference. We propose an unbiased estimator for the average treatment effect under the mixed design and show the variance of the estimator is bounded by O(d^2n^-1p^-1) where d is the maximum degree of the network, n is the network size, and p is the probability of treatment. We also establish a lower bound of Ω(d^1.5n^-1p^-1) for the variance of any mixed design. For a family of sparse networks characterized by a growth constant κ≤ d, we improve the upper bound to O(κ^7 dn^-1p^-1). Furthermore, when interference weights on the edges of the network are unknown, we propose a weight-invariant design that achieves a variance bound of O(d^3n^-1p^-1).
Title: Constant Regret Resolving Heuristics for Price-Based Revenue Management Network revenue management (NRM) and its corresponding pricing question is one of the most fundamental problems in operations management. To alleviate the curse of dimensionality and the prohibitive cost of computing an exact solution using dynamic programming, computationally efficient resolving algorithms are proposed. The state-of-the-art analysis of the resolving heuristic establishes a logarithmic additive regret for price-based NRM problems. In “Constant Regret Resolving Heuristics for Price-Based Revenue Management,” Y. Wang and H. Wang from the University of Florida and Georgia Institute of Technology, respectively, significantly advance the state-of-the-art analysis by showing a constant regret for resolving heuristics. Their theoretical advance is made possible by a novel, direct analysis of the exact DP solution.
The presence of inventory constraints is prevalent in revenue management applications and affects how pricing should be managed. This chapter reviews recent developments for the joint learning and pricing problem with inventory constraints using both frequentist and Bayesian approaches. As the total demand and supply in the system scales proportionally, information-theoretical lower bounds indicate that any algorithm must have a regret (i.e., the cumulative expected revenue loss comparing to the full-information optimal solution) that is at least in the square root order of the scaling factor. We introduce effective heuristics that match the square root regret up to multiplicative logarithmic terms. For the frequentist approach, if there is a single product, a shrinking price interval heuristic achieves square root regret. When there are multiple products, a self-adjusting heuristic achieves square root regret when the demand comes from a known class of parametric functions; if the class of functions is unknown but the demand function is sufficiently smooth, then such heuristic can attain a regret which is arbitrarily close to square root. For the Bayesian approach, a Thompson sampling-based heuristic can achieve square root regret.
Crowdsourced delivery platforms face the unique challenge of meeting dynamic customer demand using couriers not employed by the platform. As a result, the delivery capacity of the platform is uncertain. To reduce the uncertainty, the platform can offer a reward to couriers that agree to be available to make deliveries for a specified period of time, that is, to become scheduled couriers. We consider a scheduling problem that arises in such an environment, that is, in which a mix of scheduled and ad hoc couriers serves dynamically arriving pickup and delivery orders. The platform seeks a set of shifts for scheduled couriers so as to minimize total courier payments and penalty costs for expired orders. We present a prescriptive machine learning method that combines simulation optimization for off-line training and a neural network for online solution prescription. In computational experiments using real-world data provided by a crowdsourced delivery platform, our prescriptive machine learning method achieves solution quality that is within 0.2%–1.9% of a bespoke sample average approximation method while being several orders of magnitude faster in terms of online solution generation. History: This paper has been accepted for the Transportation Science Special Issue on Emerging Topics in Transportation Science and Logistics. Funding: This work was supported in part by the National Science Foundation [Grant 2145661].
We present a data-driven optimization framework for redesigning police patrol zones in an urban environment. The objectives are to rebalance police workload along geographical areas and to reduce response time to emergency calls. We develop a stochastic model for police emergency response by integrating multiple data sources, including police incident reports, demographic surveys, and traffic data. Using this stochastic model, we optimize zone-redesign plans using mixed-integer linear programming. Our proposed design was implemented by the Atlanta Police Department in March 2019. By analyzing data before and after the zone redesign, we show that the new design has reduced the response time to high-priority 911 calls by 5.8% and the imbalance of police workload among Atlanta’s zones by 43%.
We consider a freight platform that serves as an intermediary between shippers and carriers in a truckload transportation network. The platform's objective is to design a mechanism that determines prices for shippers and payments to carriers, as well as how carriers are matched to loads to be transported, to maximize its long-run average profit. We analyze three types of carrier-side mechanisms commonly used by freight platforms: posted price, auction, and a hybrid mechanism where carriers can either book loads at posted prices or submit their bids. The proposed mechanisms are constructed using a fluid approximation model to incorporate carrier interactions in the freight network. We show that the auction mechanism has higher expected profits than the hybrid mechanism, which in turn has higher profits than the posted price mechanism. Thus, the hybrid mechanism achieves a trade-off between platform profit and carrier waiting time. We prove tight bounds between these mechanisms for varying market sizes. The findings are validated through a numerical simulation using industry data from the U.S. freight market.
We consider how to allocate inventory of seasonal goods in a two-echelon distribution network for Dillard’s Inc., a large department store chain in the United States. Our objective is to allocate products with limited inventory from a distribution center to multiple retail stores over the selling season to maximize total sales revenue. Under the assumption that the true demand distributions are available to the retailer, we develop an effective dynamic inventory allocation heuristic. We further consider a more realistic and challenging setting for seasonal goods, where demand distributions are unknown to the retailer, and propose two “learning-while-doing” extensions of our inventory allocation heuristic; these policies update demand distribution estimates in a rolling horizon using censored point-of-sales data. We evaluate the performance of the policies using simulation on Dillard’s historical sales data. Dillard’s Inc. has incorporated the proposed policy into their current replenishment methodology and has been using the policy to set order levels for its seasonal merchandise.
We study a multi‐period inventory allocation problem in a one‐warehouse multiple‐retailer setting with lost sales. At the start of a finite selling season, a fixed amount of inventory is available at the warehouse. Inventory can be allocated to the retailers over the course of the selling horizon (transshipment is not allowed). The objective is to minimize the total expected lost sales and holding costs. In each period, the decision maker can use the realized and possibly censored demand observations to dynamically update demand forecast and consequently make allocation decisions. Our model allows a general demand updating framework, which includes ARMA models or Bayesian methods as special cases. We propose a computationally tractable algorithm to solve the inventory allocation problem under demand learning using a Lagrangian relaxation technique, and show that the algorithm is asymptotically optimal. We further use this technique to investigate how demand learning would affect inventory allocation decisions in a two‐period setting. Using a combination of theoretical and numerical analysis, we show that demand learning provides an incentive for the decision maker to withhold inventory at the warehouse rather than allocating it in early periods.
We consider a two-sided marketplace in which a market operator sells services to clients and buys services from vendors. The market operator determines the prices for both clients and vendors. Services are transacted in discrete units called jobs, and the jobs are characterized by job types, service deadlines, payments from clients, and payments to vendors. Jobs are submitted by clients to the marketplace, and the market operator then lists the available jobs. Vendors can view the available jobs with their job characteristics, and choose jobs based on their preferences. We consider an infinite horizon long-run average reward Markov decision process (MDP) model of the market operator's pricing problem. The MDP includes random arrival processes on both sides of the marketplace, as well as the choice behavior of clients and vendors. Since solving this MDP directly is impractical for large market sizes, we study a discrete-time fluid approximation of the problem. This approximation results in a simple pricing policy whose prices for each job depend on the time remaining until the service deadline, but do not depend on the other jobs available in the marketplace. We show that this policy is asymptotically optimal with a loss ratio of order $O(1/\theta)$ on the long-run average reward, where $\theta$ represents the scale of demand and supply. We also present a continuous-time fluid model and discuss the managerial insights that follow from its solution.
We consider an integrated pricing and routing problem on a service network motivated by environments encountered at package express carriers. The decision maker sets the price for each origin-destination market, which determines the demand that needs to be served. The demand can be routed along multiple paths in the service network if desirable. The objective is to maximize the revenues from serving demand minus the transportation costs incurred by serving demand given the capacities in the network. We propose two algorithms for the solution of this problem with theoretical convergence guarantees: (1) a Frank-Wolfe type algorithm, which requires the objective function to be smooth, and (2) a primal-dual algorithm using an online learning technique, which allows non-smooth objective functions. We show that both algorithms have a convergence rate of Õ(1/T ) where T is the number of iterations. Numerical experiments on randomly generated instances show that coordinating pricing and routing decisions can improve profits by more than 10%.
We consider a canonical quantity-based network revenue management problem where a firm accepts or rejects incoming customer requests irrevocably in order to maximize expected revenue given limited resources. Because of the curse of dimensionality, the exact solution to this problem by dynamic programming is intractable when the number of resources is large. We study a family of re-solving heuristics that periodically re-optimize an approximation to the original problem known as the deterministic linear program (DLP), where random customer arrivals are replaced by their expectations. We find that, in general, frequently re-solving the DLP produces the same order of revenue loss as one would get without re-solving, which scales as the square root of the time horizon length and resource capacities. By re-solving the DLP at a few selected points in time and applying thresholds to the customer acceptance probabilities, we design a new re-solving heuristic with revenue loss that is uniformly bounded by a constant that is independent of the time horizon and resource capacities. This paper was accepted by Kalyan Talluri, revenue management and market analytics.
We redesign the police patrol beat in South Fulton, Georgia, in collaboration with the South Fulton Police Department (SFPD), using a predictive data-driven optimization approach. Due to rapid urban development and population growth, the existing police beat design done in the 1970s was far from efficient, which leads to low policing efficiency and long 911 call response time. We balance the police workload among different city regions, improve operational efficiency, and reduce 911 call response time by redesigning beat boundaries for the SFPD. We discretize the city into small geographical atoms, which correspond to our decision variables; the decision is to map the atoms into "beats", the basic unit of the police operation. We first analyze workload and trend in each atom using the rich dataset, including police incidents reports and U.S. census data; We then predict future police workload for each atom using spatial statistical regression models; Lastly, we formulate the optimal beat design as a mixed-integer programming (MIP) program with continuity and compactness constraints on the beats' shape. The optimization problem is solved using simulated annealing due to its large-scale and non-convex nature. The simulation results suggest that our proposed beat design can reduce workload variance among beats significantly by over 90\%.
We study a multiperiod dynamic pricing problem with contextual information, where the seller uses a misspecified demand model. The seller sequentially observes past demand, updates model parameters, and then chooses the price for the next period based on time-varying features. We show that model misspecification leads to a correlation between price and prediction error of demand per period, which, in turn, leads to inconsistent price elasticity estimates and hence suboptimal pricing decisions. We propose a “random price shock” (RPS) algorithm that dynamically generates randomized price shocks to estimate price elasticity, while maximizing revenue. We show that the RPS algorithm has strong theoretical performance guarantees, that it is robust to model misspecification, and that it can be adapted to a number of business settings, including (1) when the feasible price set is a price ladder and (2) when the contextual information is not IID. We also perform offline simulations to gauge the performance of RPS on a large fashion retail data set and find that is expected to earn 8%–20% more revenue on average than competing algorithms that do not account for price endogeneity. This paper was accepted by Serguei Netessine, operations management.
We consider Markov Decision Processes (MDPs) where the rewards are unknown and may change in an adversarial manner. We provide an algorithm that achieves state-of-the-art regret bound of $O( \sqrt{\tau (\ln|S|+\ln|A|)T}\ln(T))$, where $S$ is the state space, $A$ is the action space, $\tau$ is the mixing time of the MDP, and $T$ is the number of periods. The algorithm's computational complexity is polynomial in $|S|$ and $|A|$ per period. We then consider a setting often encountered in practice, where the state space of the MDP is too large to allow for exact solutions. By approximating the state-action occupancy measures with a linear architecture of dimension $d\ll|S|$, we propose a modified algorithm with computational complexity polynomial in $d$. We also prove a regret bound for this modified algorithm, which to the best of our knowledge this is the first $\tilde{O}(\sqrt{T})$ regret bound for large scale MDPs with changing rewards.