
This paper studies a variant of two-player zero-sum matrix games, where, at each timestep, the row player selects row i, the column player selects column j, and the row player receives a noisy reward with expected value A_i,j , along with noisy feedback on the input matrix A . The row player’s goal is to maximize their total reward against an adversarial column player. Nash regret, defined as the difference between the player’s total reward and the game’s Nash equilibrium value scaled by the time horizon T, is often used to evaluate algorithmic performance in zero-sum games . We begin by studying the limitations of existing algorithms for minimizing Nash regret in zero-sum games. We show that standard algorithms—including Hedge, FTRL, and OMD—as well as the strategy of playing the Nash equilibrium of the empirical matrix—all incur (√(T)) Nash regret, even when the row player receives noisy feedback on the entire matrix A. Furthermore, we show that UCB for matrix games, a natural adaptation of the well-known bandit algorithm, also suffers (√(T)) Nash regret under bandit feedback. Notably, these lower bounds hold even in the simplest case of 2 × 2 matrix games, where the instance-dependent matrix parameters are constant. This highlights a fundamental limitation of existing methods: they fail to leverage favorable problem structure to achieve lower regret. Motivated by this limitation, we ask whether instance-dependent polylog(T) Nash regret is achievable against adversarial opponents. We answer this affirmatively. In the full-information setting, we present the first algorithm for general n × m matrix games that achieves instance-dependent polylog(T) Nash regret. In the bandit feedback setting, we design an algorithm with similar guarantees for the special case of 2 × 2 games—the same regime in which existing algorithms provably suffer (√(T)) regret despite the simplicity of the instance. Finally, we validate our theoretical results with empirical evidence.
In job-scheduling games, each job is a selfish player that selects a machine to minimize its own completion time. Coordination mechanisms are employed to reduce the inefficiency of equilibria that result from such decentralized decision-making. This paper contributes to the extensive body of research on coordination mechanisms by investigating their application to unrelated parallel machines, where each machine may use its own scheduling policy to determine the processing order of assigned jobs. Since pure Nash equilibria (NE) are not guaranteed to exist in this setting, we identify and characterize several classes of instances-motivated by real-world applications-in which a NE is guaranteed to exist. For each such class, we design an algorithm to compute a NE, prove the convergence of best-response dynamics, and analyze the inefficiency of equilibria with respect to the makespan. In addition, we study two fundamental problems: (1) computing a NE schedule with low makespan, and (2) selecting, given a matrix of processing times, machine-specific scheduling policies that guarantee the existence of a NE with low makespan. For both problems, we establish computational hardness results.
Through the lens of Tullock contests, we take a fresh look at the contest game for crowdsourcing reviews [11] as a Tullock contest game with discrete strategies: Each of n players, endowed with skills, strategically invests an effort for writing a review of some quality she chooses from a finite set [Q]; she is awarded a payment, out of a budget β , in proportion to her effort, and pays a player-specific cost, which is the product of effort and skill. So both the payment and the utility functions for a player are strictly concave in the player’s effort. Players are anonymous if skills are identical, and so is the game. We study the mixed Nash equilibria of the game, where no player could deviate to increase her expected utility; her support in a Nash equilibrium is the set of qualities she chooses with strictly positive probability. By strict concavity, mixed Nash equilibria have the Small Consecutive-Supports property: their supports have size at most 2. We show:
We provide the first forecasting competition mechanism that is both truthful and accurate for arbitrarily correlated events. Previous work describes mechanisms that either work only for independent events or block-correlated events or that do not ensure accuracy. Our mechanism works for any event structure that can be represented as a decomposable Bayesian network, an assumption that is without loss of generality, since any joint distribution can be represented as a Bayes net (BN), and any BN can be made decomposable. We generalize the Event-Lotteries Forecaster Selection Mechanism (ELF) which works for independent events [12,13]. We first show that ELF is truthful for any two events, regardless of correlation. We next show that the counterexample of Witkowski et al. [12] for correlated events is circumvented if forecasters can update their beliefs over time. We then create a new counterexample of only three correlated events showing that ELF fails even if forecasters can update. Finally, we design a new competition mechanism called Bayesian network ELF (BNELF) that is truthful for any structure of correlation represented as a decomposable BN (e.g., a tree or join tree). BNELF divides the competition into two subgames. By asking forecasters to provide probabilities for each event E in one subgame conditioned on all possible combinations of outcomes of E's children in the BN in the second subgame, the events in each subgame become conditionally independent. Finally, BNELF flips a coin; forecasters play ELF randomly in one of the two subgames. We show that, given enough events that are at least minimally uncertain after conditioning, BNELF is guaranteed to accurately converge to identify the best forecaster.
We consider a variation of the classic Hotelling-Downs model with the addition of facility synergies. Unlike in the classic model, where clients always use the facility closest to them, we study clients who prefer locations with many facilities to those with few facilities while simultaneously attempting to minimize their distance as well. We show that, in contrast with the classic model, Nash equilibria for our setting always exist, and, in fact, there always exists a Nash equilibrium such that the sum of client costs equals the cost of the optimum solution. Our main result is a bound of 225/64≈ 3.516 on the Price of Anarchy for our model, showing that, although the client behavior is more complex in our model (and often more realistic depending on the application), the cost of Nash equilibrium solutions still cannot be much worse than the cost of the optimum facility placement.
Prominent opinion formation models such as the one by Friedkin and Johnsen (FJ) concentrate on the effects of peer pressure on public opinions. In practice, opinion formation is also based on information about the state of the world and persuasion efforts. In this paper, we analyze an approach of Bayesian persuasion in the FJ model. There is an unknown state of the world that influences the preconceptions of n agents. A sender 𝒮 can (partially) reveal information about the state to all agents. The agents update their preconceptions, and an equilibrium of public opinions emerges. We propose algorithms for the sender to reveal information in order to optimize various aspects of the emerging equilibrium. For many natural sender objectives, we show that there are simple optimal strategies. We then focus on a general class of range-based objectives based on desired opinion ranges for each agent. We provide efficient algorithms in several cases, e.g., when the matrix of preconceptions in all states has constant rank, or when there is only a polynomial number of range combinations that lead to positive value for 𝒮 . This generalizes, e.g., instances with a constant number of states and/or agents, or instances with a logarithmic number of ranges. In general, we show that subadditive range-based objectives allow a simple n-approximation, and even for additive ones, obtaining an n^1-ε -approximation is NP-hard, for any constant ε > 0 .
The recent rise of renewable energy produced by many decentralized sources yields interesting market design challenges for electrical grids. Balancing supply and demand in such networks is both a temporal and spatial challenge due to capacity constraints. The recent surge in the number of household-owned batteries, especially in regions with rooftop solar adoption, offers mitigation potential but often acts misaligned with grid-level objectives. In fact, the decision to charge or discharge a household-owned battery is a strategic choice by each battery owner governed by selfish incentives. This calls for an analysis from a game-theoretic point of view. We initiate this timely research direction by considering a game-theoretic setting where selfish agents strategically charge or discharge their batteries to increase their profit. In particular, we study a Stackelberg-like market model where a third party introduces price incentives, aiming to optimize renewable energy utilization while preserving grid feasibility. For this, we study the existence and the quality of equilibria under various pricing strategies. We find that the existence of equilibria crucially depends on the chosen pricing and that the obtained social welfare varies widely. This calls for more sophisticated market models and pricing mechanisms and opens up a rich field for future research in Algorithmic Game Theory on incentives in renewable energy networks.
We study the fair allocation of indivisible goods under cardinality constraints, where each agent must receive a bundle of fixed size. This models practical scenarios, such as assigning shifts or forming equally sized teams. Recently, variants of envy-freeness up to one/any item (EF1, EFX) were introduced for this setting, based on flips or exchanges of items. Namely, one can define envy-freeness up to one/any flip (EFF1, EFFX), meaning that an agent i does not envy another agent j after performing one or any one-item flip between their bundles that improves the value of i. We explore algorithmic aspects of this notion, and our contribution is twofold: we present both algorithmic and impossibility results, highlighting a stark contrast between the classic EFX concept and its flip-based analogue. First, we explore standard techniques used in the literature and show that they fail to guarantee EFFX approximations. On the positive side, we show that we can achieve a constant factor approximation guarantee when agents share a common ranking over item values, based on the well-known envy cycle elimination technique. This idea also leads to a generalized algorithm with approximation guarantees when agents agree on the top n items and their valuation functions are bounded. Finally, we show that an algorithm that maximizes the Nash welfare guarantees a 1/2-EFF1 allocation, and that this bound is tight.
We study deterministic mechanisms for the two-facility location problem. Given the reported locations of n agents on the real line, such a mechanism specifies where to build the two facilities. The single-facility variant of this problem admits a simple strategyproof mechanism that minimizes social cost. For two facilities, however, it is known that any strategyproof mechanism is (n) -approximate. We seek to circumvent this strong lower bound by relaxing the problem requirements. Following other work in the facility location literature, we consider a relaxed form of strategyproofness in which no agent can lie and improve their outcome by more than a constant factor. Because the aforementioned (n) lower bound generalizes easily to constant-strategyproof mechanisms, we introduce a second relaxation: Allowing the facilities (but not the agents) to be located in the plane. Our first main result is a natural mechanism for this relaxation that is constant-approximate and constant-strategyproof. A characteristic of this mechanism is that a small change in the input profile can produce a large change in the solution. Motivated by this observation, and also by results in the facility re-allocation literature, our second main result is a constant-approximate, constant-strategyproof, and Lipschitz continuous mechanism.
We consider a mechanism design setting with a single item and a single buyer who is uncertain about the value of the item. Both the buyer and the seller have a common model for the buyer’s value, but the buyer discovers her true value only upon receiving the item. Mechanisms in this setting can be interpreted as randomized refund mechanisms, which allocate the item at some price and then offer a (partial and/or randomized) refund to the buyer in exchange for the item if the buyer is unsatisfied with her purchase. Motivated by their practical importance, we study the design of optimal deterministic mechanisms in this setting. We characterize optimal mechanisms as virtual value maximizers for both continuous and discrete type settings. We then use this characterization, along with bounds on the menu size complexity, to develop efficient algorithms for finding optimal and near-optimal deterministic mechanisms (The full paper can be accessed at https://arxiv.org/abs/2507.04148 ).
We consider an optimal stopping problem with n correlated offers where the goal is to design a (randomised) stopping strategy that maximises the expected value of the offer at which we stop. Instead of assuming to know the complete correlation structure, which is unrealistic in practice, we only assume to have knowledge of the distribution of the maximum value X-max of the sequence, and want to analyse the worst-case correlation structure whose maximum follows this distribution. This can be seen as a trade-off between the setting in which no distributional information is known, and the Bayesian setting in which the (possibly correlated) distributions of all the individual offers are known. As our first main result we show that a deterministic threshold strategy using the monopoly price of the distribution of the maximum value is asymptotically optimal assuming that the expectation of the maximum value grows sublinearly in n. In our second main result, we further tighten this bound by deriving a tight quadratic convergence guarantee for sufficiently smooth distributions of the maximum value. Our results also give rise to a more fine-grained picture regarding prophet inequalities with correlated values, for which distribution-free bounds only yield a performance guarantee of the order 1/n.
Since its introduction, envy-freeness up to any good (EFX) has become a fundamental solution concept in fair division of indivisible goods. Its existence, however, remains elusive-even for four agents with additive utility functions, it is unknown whether an EFX allocation always exists. Unsurprisingly, researchers have explored restricted settings to delineate tractable and intractable cases. Christadolou, Fiat et al. [EC'23] introduced the notion of EFX-orientation, where the agents form the vertices of a graph and the items correspond to edges, and an agent values only the items that are incident to it. The goal is to allocate items to one of the adjacent agents while satisfying the EFX condition. This graph based setting has received considerable attention and has led to a growing body of work. Building on the work of Zeng and Mehta'24, which established a sharp complexity threshold based on the structure of the underlying graph-polynomial-time solvability for bipartite graphs and NP-hardness for graphs with chromatic number at least three-we further explore the algorithmic landscape of EFX-orientation by exploiting the graph structure using parameterized graph algorithms. Specifically, we show that bipartiteness is a surprisingly stringent condition for tractability: EFX orientation is NP-complete even when the valuations are symmetric, binary and the graph is at most two edge-removals away from being bipartite. Moreover, introducing a single non-binary value makes the problem NP-hard even when the graph is only one edge removal away from being bipartite. We further perform a parameterized analysis to examine structures of the underlying graph that enable tractability. In particular, we show that the problem is solvable in linear time on graphs whose treewidth is bounded by a constant. Furthermore, we also show that the complexity of an instance is closely tied to the sizes of acyclic connected components on its one-valued edges.
Augmenting the input of algorithms with predictions is an algorithm design paradigm that suggests leveraging a (possibly erroneous) prediction to improve worst-case performance guarantees when the prediction is perfect (consistency), while also providing a performance guarantee when the prediction fails (robustness). Recently, Xu at al. [40] and Agrawal et al. [1] proposed to consider settings with strategic agents under this framework. In this paper, we initiate the study of budget-feasible mechanism design with predictions. These mechanisms model a procurement auction scenario in which an auctioneer (buyer) with a strict budget constraint seeks to purchase goods or services from a set of strategic agents, so as to maximize her own valuation function. We focus on the online version of the problem where the arrival order of agents is random. We design mechanisms that are truthful, budget-feasible, and achieve a significantly improved competitive ratio for both monotone and non-monotone submodular valuation functions compared to their state-of-the-art counterparts without predictions. Our results assume access to a prediction for the value of the optimal solution to the offline problem. We complement our positive results by showing that for the offline version of the problem, access to predictions is mostly ineffective in improving approximation guarantees.
We study an online fair division setting, where goods arrive one at a time and there is a fixed set of n agents, each of whom has an additive valuation function over the goods. Once a good appears, the value each agent has for it is revealed and it must be allocated immediately and irrevocably to one of the agents. It is known that without any assumptions about the values being severely restricted or coming from a distribution, very strong impossibility results hold in this setting [21, 28]. To bypass the latter, we turn our attention to instances where the valuation functions are restricted. In particular, we study personalized 2-value instances, where there are only two possible values each agent may have for each good, possibly different across agents, and we show how to obtain worst case guarantees with respect to well-known fairness notions, such as maximin share fairness and envy-freeness up to one (or two) good(s). We suggest a deterministic algorithm that maintains a 1/(2n-1) - MMS allocation at every time step and show that this is the best possible any deterministic algorithm can achieve if one cares about every single time step; nevertheless, eventually the allocation constructed by our algorithm becomes a 1/4- MMS allocation. To achieve this, the algorithm implicitly maintains a fragile system of priority levels for all agents. Further, we show that, by allowing some limited access to future information, it is possible to have stronger results with less involved approaches. In particular, by knowing the values of goods for n-1 time steps into the future, we design a matching-based algorithm that achieves an EF1 allocation every n time steps, while always maintaining an EF2 allocation.
In the Course Allocation problem, there are a set of students and a set of courses at a given university. University courses may have different numbers of credits, typically related to different numbers of learning hours, and there may be other constraints such as courses running concurrently. Our goal is to allocate the students to the courses such that the resulting matching is stable, which means that no student and course(s) have an incentive to break away from the matching and become assigned to one another. We study several definitions of stability and for each we give a mixture of polynomial-time algorithms and hardness results for problems involving verifying the stability of a matching, finding a stable matching or determining that none exists, and finding a maximum size stable matching. We also study variants of the problem with master lists of students, and lower quotas on the number of students allocated to a course, establishing additional complexity results in these settings.
We study the Stable Fixtures problem, a many-to-many generalisation of the classical non-bipartite Stable Roommates matching problem. Building on the foundational work of Tan on stable partitions, we extend his results to this significantly more general setting and develop a rich framework for understanding stable structures in many-to-many contexts. Our main contribution, the notion of a generalised stable partition (GSP), not only characterises the solution space of this problem, but also serves as a versatile tool for reasoning about ordinal preference systems with capacity constraints. We show that a GSP can be computed efficiently and can provide an elegant representation of key aspects of a preference system. Leveraging a connection to stable half-matchings, we also establish a non-bipartite analogue of the Rural Hospitals Theorem for stable half-matchings and GSPs, and connect our results to recent work on near-feasible matchings, providing a simpler algorithm and tighter analysis for this problem. Beyond structural insights, we conduct the first empirical analysis of random Stable Fixtures instances, uncovering surprising results, such as the impact of capacity functions on the solvability likelihood.
We study the problem of fairly and efficiently allocating a set of items among strategic agents with additive valuations, where items are either all indivisible or all divisible. When items are goods, numerous positive and negative results are known regarding the fairness and efficiency guarantees achievable by truthful mechanisms, whereas our understanding of truthful mechanisms for chores remains considerably more limited. In this paper, we discover various connections between truthful good and chore allocations, greatly enhancing our understanding of the latter via tools from the former. For indivisible chores with two agents, by leveraging the observation that a simple bundle-swapping operation transforms several properties for goods including truthfulness to the corresponding properties for chores, we characterize truthful mechanisms and derive tight guarantees of various fairness notions achieved by truthful mechanisms. Moreover, for homogeneous divisible chores, by generalizing the above transformation to an arbitrary number of agents, we characterize truthful mechanisms with two agents, show that every truthful mechanism with two agents admits an efficiency ratio of 0, and derive a large family of strictly truthful, envy-free (EF), and proportional mechanisms for an arbitrary number of agents. Finally, for indivisible chores with an arbitrary number of agents having bi-valued cost functions, we give an ex-ante truthful, ex-ante Pareto optimal, ex-ante EF, and ex-post envy-free up to one item mechanism, improving the best guarantees for bi-valued instances by prior works.
This paper examines the impact of agents' myopic optimization on the efficiency of systems comprised by many selfish agents. In contrast to standard congestion games where agents interact in a one-shot fashion, in our model each agent chooses an infinite sequence of actions and maximizes the total reward stream discounted over time under different ways of computing present values. Our model assumes that actions consume common resources that get congested, and the action choice by an agent affects the completion times of actions chosen by other agents, which in turn affects the time rewards are accrued and their discounted value. This is a mean-field game, where an agent's reward depends on the decisions of the other agents through the resulting action completion times. For this type of game we define stationary equilibria, and analyze their existence and price of anarchy (PoA). Overall, we find that the PoA depends entirely on the type of discounting rather than its specific parameters. For exponential discounting, myopic behaviour leads to extreme inefficiency: the PoA is infinity for any value of the discount parameter. For power law discounting, such inefficiency is greatly reduced and the PoA is 2 whenever stationary equilibria exist. This matches the PoA when there is no discounting and players maximize long-run average rewards. Additionally, we observe that exponential discounting may introduce unstable equilibria in learning algorithms, if action completion times are interdependent. In contrast, under no discounting all equilibria are stable.
With the rise of online applications, recommender systems (RSs) often encounter constraints in balancing exploration and exploitation. Such constraints arise when exploration is carried out by agents whose utility must be taken into account when optimizing overall welfare. Recent work suggests that recommendations should be mechanism-informed individually rational (MIR) [6]. Specifically, if agents have a default arm they would use, relying on the RS should yield each agent at least the reward of the default arm, conditioned on the information available to the RS. Under the MIR constraint, striking a balance between exploration and exploitation becomes a complex planning problem. To that end, Bahar et al. [6] propose an approximately optimal yet inefficient planning algorithm that runs in O(2(K) K-2 H-2), where K is the number of arms and H is the size of the support of the reward distributions. In this paper, we make a significant improvement for a special yet practical case, removing both the dependence on.H and the exponential dependence on K. We assume a stochastic order of the rewards (e.g., Gaussian with unit variance, Bernoulli, etc.), and devise an asymptotically optimal algorithm with a runtime of O(K logK). Our technique is based on formulating a Goal Markov Decision Process (GMDP), establishing an optimal dynamic programming procedure, and then unveiling its crux-fleshing out a simple index-based structure that facilitates efficient computation. Additionally, we present an incentive-compatible version of our algorithm.
A classical problem in combinatorics seeks colorings of low discrepancy. More concretely, the goal is to color the elements of a set system so that the number of appearances of any color among the elements in each set is as balanced as possible. We present a new lower bound for multicolor discrepancy, showing that there is a set system with n subsets over a set of elements in which any k-coloring of the elements has discrepancy at least ( √(n/lnk)) . This result improves the previously best-known lower bound of ( √(n/k)) of Doerr and Srivastav [10] and may have several applications. Here, we explore its implications on the feasibility of fair division concepts for instances with n agents having valuations for a set of indivisible items. The first such concept is known as consensus 1/k-division up to d items (CDd) and aims to allocate the items into k bundles so that no matter which bundle each agent is assigned to, the allocation is envy-free up to d items. The above lower bound implies that CDd can be infeasible for d∈( √(n/lnk)) . We furthermore extend our proof technique to show that there exist instances of the problem of allocating indivisible items to k groups of n agents in total so that envy-freeness and proportionality up to d items are infeasible for d∈( √(n/klnk)) and d∈( √(n/k^3lnk)) , respectively. The lower bounds for fair division improve the currently best-known ones by Manurangsi and Suksompong [16].