Many multi-agent socio-technical systems rely on aggregating heterogeneous agents' costs into a social cost function (SCF) to coordinate resource allocation in domains such as energy grids, water allocation, or traffic management. The choice of SCF often entails implicit assumptions and may lead to undesirable outcomes if not rigorously justified. In this paper, we demonstrate that what determines which SCF ought to be used is the degree to which individual costs can be compared across agents and which axioms the aggregation shall fulfill. Drawing on the results from social choice theory, we provide guidance on how this process can be used in control applications. We demonstrate which assumptions about interpersonal utility comparability- ranging from ordinal level comparability to full cardinal comparability- together with a choice of desirable axioms, inform the selection of a correct SCF, be it the classical utilitarian sum, the Nash SCF, or maximin. Thus, fixing comparability level first, then choosing an objective from the compatible class, and reporting both as part of the specification, makes the fairness and efficiency consequences transparent. We demonstrate how the proposed framework can be applied for principled allocations of water, transportation, and energy resources.
Digital marketplaces processing billions of dollars annually represent critical infrastructure in sociotechnical ecosystems, yet their performance optimization lacks principled measurement frameworks that can inform algorithmic governance decisions regarding market efficiency and fairness from complex market data. By looking at orderbook data from double auction markets alone, because bids and asks do not represent true maximum willingnesses to buy and true minimum willingnesses to sell, there is little an economist can say about the market’s actual performance in terms of allocative efficiency. We turn to experimental data to address this issue, ‘inverting’ the standard induced value approach of double auction experiments. Our aim is to predict key market features relevant to market efficiency, particularly allocative efficiency, using orderbook data only—specifically bids, asks and price realizations, but not the induced reservation values—as early as possible. Since there is no established model of strategically optimal behavior in these markets, and because orderbook data is highly unstructured, non-stationary and non-linear, we propose quantile-based normalization techniques that help us build general predictive models. We develop and train several models, including linear regressions and gradient boosting trees, leveraging quantile-based input from the underlying supply-demand model. Our models can predict allocative efficiency with reasonable accuracy from the earliest bids and asks, and these predictions improve with additional realized price data. The performance of the prediction techniques varies by target and market type. Our framework holds significant potential for application to real-world market data, offering valuable insights into market efficiency and performance, even prior to any trade realizations.
Information in repeated real-world interactions is rarely fixed. Instead, over time, players typically learn more about the game structure, understand more about their own payoff correspondences, and find out what others did and earned in the past. However, how this dynamic information path influences learning and which aggregate outcomes are reached as a result have not been investigated to date. To study this, we conducted a series of laboratory experimental games where we provided more information over time in different orders and along different paths. These games span two strategically distinct classes: a Cournot game with a unique symmetric Nash equilibrium, and step-level public-goods (coordination) games with multiple Nash equilibria of differing Pareto efficiency. Our evidence confirms a natural mapping from information to predominant learning rule: information about own payoffs triggers payoff-based learning, feedback concerning others’ realized payoffs triggers imitation, and structural information about the game triggers best response. In addition, when multiple learning rules are feasible, which one dominates depends on information paths. In particular, feedback about others’ actions and payoffs, especially when supplied marginally, may trigger persistent imitation that locks into Pareto-inferior Nash equilibria, even when more information later becomes available.
Minimizing volatility and adjustment costs is of central importance in many economic environments, yet it is often complicated by evolving feasibility constraints. We study a decision maker who repeatedly selects an action from a stochastically evolving interval of feasible actions in order to minimize either average adjustment costs or variance. We show that for strictly convex adjustment costs (such as quadratic variation), the optimal decision rule is a reference rule in which the decision maker minimizes the distance to a target action. In general, the optimal target depends both on the previous action and the expectation of future constraints; but for the special case where the constraints follow a random walk, the optimal mechanism is to simply target the previous action. If the decision maker minimizes variance, the optimal policy is also a reference rule, but the target is a constant, which is not necessarily equal to the long-term average action. Compared to mid-point heuristics, these optimal rules may substantially reduce quadratic variation and variance, in natural environments by 50% or more. Applied to stock market auctions, our results provide an explanation for the wide-spread use of reference price rules. We also apply our results to bilateral trade in over-the-counter markets, capacity planning in supply chains, and positioning in political agenda setting.
Pursuing replicability - independent evidence for previous claims - is important for creating generalizable knowledge(1,2). Here we attempted replications of 274 claims of positive results from 164 quantitative papers published from 2009 to 2018 in 54 journals in the social and behavioural sciences. Replications were high powered on average to detect the original effect size (median of 99.6%), used original materials when relevant and available, and were peer reviewed in advance through a standardized internal protocol. Replications showed statistically significant results in the original pattern for 151 of 274 claims (55.1% (95% confidence interval (CI) 49.2-60.9%)) and for 80.8 of 164 papers (49.3% (95% CI 43.8-54.7%)), weighed for replicating multiple claims per paper. We observed modest variation in replication rates across disciplines (42.5-63.1%), although some estimates had high uncertainty. The median Pearson's r effect size was 0.25 (95% CI 0.21-0.27) for original studies and 0.10 (95% CI 0.09-0.13) for replication studies, an 82.4% (95% CI 67.8-88.2%) reduction in shared variance. Thirteen methods for evaluating replication success provided estimates ranging from 28.6% to 74.8% (median of 49.3%). Some decline in effect size and significance is expected based on power to detect original effects and regression to the mean because we replicated only positive results. We observe that challenges for replicability extend across social-behavioural sciences, illustrating the importance of identifying conditions that promote or inhibit replicability(3,4).
At the core of most socio-technical systems lies a scarce resource that is allocated among agents: highway lanes, public transit, road space, water rights, energy access, grid capacity, user attention, pollution rights, etc. With further automation of the underlying allocation processes, control engineers are increasingly tasked to make decisive assumptions regarding what society wants. In practice to date, design choices are largely driven by industry norms and conventions rather than a result of conscientiously responsible and ethical design. In this paper, we look at tools available to control engineers to design systems in a more principled manner in order to match the societal mandate. We consider three control design paradigms: online feedback optimization, control of Markov decision processes, and model predictive control. Beginning with aggregating individual agents' preferences into control design objectives, subsequently ensuring and certifying the fulfillment of those specifications, we argue that the feedback nature of control systems enables appropriate allocation of the shared resources in ways hitherto unparalleled.
The Price of Anarchy (PoA) is a popular measure of the costs of decentralization in terms of efficiency losses. Almost all PoA analyses operate within a framework assuming both Cardinal Full-Comparability (CFC) and smoothness, in which case any derived bounds conveniently extend beyond pure Nash to coarse correlated equilibria and no-regret learning outcomes. However, interpersonal utility comparability is an additional assumption that generally has to be justified. Without it, cardinal utilities (e.g. defined under classical von Neumann–Morgenstern framework) are unique only up to agent-specific affine transformations, rendering both the utilitarian PoA and the classical smoothness conditions representation-dependent. In this paper, we operate under a more general Cardinal Non-Comparability (CNC) framework, under which the weighted Nash welfare is a canonical admissible aggregator. We introduce multiplicative smoothness, a product-form condition matched to the multiplicative structure of Nash welfare, and obtain PoA bounds that are CNC-invariant and extend to coarse correlated equilibria. We demonstrate applicability of our framework on single-choice welfare games, deriving the bounds through simple proof relying on multiplicative retention envelope and geometric closure. The interpretation of this bound in terms of the true cost of decentralization depends crucially on interpersonal comparability of utilities.
Von Neumann’s minimax theorem defines optimal strategic unpredictability in zero-sum games. Empirical evidence from professional sports has been interpreted as positive behavioral evidence for minimax. In this article, we analyze the strategic optimality of offensive plays in the basketball endgame when a team has a final possession and trails by no more than a single basket. This final moment of the game most closely approximates the simultaneous-move conditions of a game where minimax theory applies. Using comprehensive NBA data from 2010 to 2025, we test for equality of success rates across shooter types (star vs. non-stars) and shot selection (two-point vs. three-point). Our analysis reveals systematic violations of minimax play that have intensified with basketball’s shift to three-pointers and higher expected points. In the final decisive moment of the game, we find that teams systematically overuse three-point shots even though the two-point attempt yields higher field goal percentages. In addition, teams over-rely on star players for the final shot; non-star two-point shots have been the top-performing endgame option in 2022–2025.
We review the experimental learning in games literature, and identify avenues for behavioral mechanism design. The focus of this review will be on recent results that investigate the effects of feedback and information on the behavior of humans in dynamic environments involving other humans and algorithms. These dynamic environments include repeated double-sided auctions and karma auctions. We also review methods to quantify dynamic effects in experimental data, and discuss new behavioral trends that arise due to the dynamics.
Generalized Nash equilibrium (GNE) problems are commonly used to model strategic interactions between self-interested agents who are coupled in cost and constraints. Specifically, the variational GNE, a refinement of the GNE, is often selected as the solution concept due to its nondiscriminatory treatment of agents by charging a uniform "shadow price" for shared resources. We study the fairness concept of v-GNEs from a comparability perspective and show that it makes an implicit assumption of unit comparability of agent’s cost functions, one of the strongest comparability notions. Further, we introduce a new solution concept, f-GNE in which a fairness metric is chosen a priori which is compatible with the comparability at hand. We introduce an electric vehicle charging game to demonstrate the fragility of v-GNE fairness and compare it to the f-GNE under various fairness metrics.
Designing incentives for an adapting population is a ubiquitous problem in a wide array of economic applications and beyond. In this work, we study how to design additional rewards to steer multi-agent systems towards desired policies \emph{without} prior knowledge of the agents' underlying learning dynamics. Motivated by the limitation of existing works, we consider a new and general category of learning dynamics called \emph{Markovian agents}. We introduce a model-based non-episodic Reinforcement Learning (RL) formulation for our steering problem. Importantly, we focus on learning a \emph{history-dependent} steering strategy to handle the inherent model uncertainty about the agents' learning dynamics. We introduce a novel objective function to encode the desiderata of achieving a good steering outcome with reasonable cost. Theoretically, we identify conditions for the existence of steering strategies to guide agents to the desired policies. Complementing our theoretical contributions, we provide empirical algorithms to approximately solve our objective, which effectively tackles the challenge in learning history-dependent strategies. We demonstrate the efficacy of our algorithms through empirical evaluations.
A series of articles has tested von Neumann’s minimax theory against behavioral evidence based on field data from professional sports. The evidence has been viewed and collectively cited as positive evidence that elite athletes in their familiar sports contexts mix well and behave in line with minimax. In this paper, based on open state-of-the-art tennis data and analytics, we shall uncover new and significant evidence against minimax at the very top of the game, where previously, such results had not been obtained. The kinds of behavioral deviations from minimax that we find become apparent, because we enrich the test strategy to take into account whether or not players face ‘pressure’ situations like break points and other decisive points. Our paper highlights that the prior literature’s failure to reject minimax does not constitute positive behavioral evidence, as some of that literature argued, because it is not robust to data aggregations and separations that are psychologically natural given the relevant real-world context. In this case, this means separating serves into the serve types that players actually consider and separating situations by pressure levels, which leads to clear and sound rejection of minimax.
The Price of Anarchy (PoA) is a standard metric for quantifying inefficiency in socio-technical systems, widely used to guide policies like traffic tolling. Conventional PoA analysis relies on exact numerical costs. However, in many settings, costs represent agents' preferences and may be defined only up to possibly arbitrary scaling and shifting, representing informational and modeling ambiguities. We observe that while such transformations preserve equilibrium and optimal outcomes, they change the PoA value. To resolve this issue, we rely on results from Social Choice Theory and define the Invariant PoA. By connecting admissible transformations to degrees of comparability of agents' costs, we derive the specific social welfare functions which ensure that efficiency evaluations do not depend on arbitrary rescalings or translations of individual costs. Case studies on a toy example and the Zurich network demonstrate that identical tolling strategies can lead to substantially different efficiency estimates depending on the assumed comparability. Our framework thus demonstrates that explicit axiomatic foundations are necessary in order to define efficiency metrics and to appropriately guide policy in large-scale infrastructure design robustly and effectively.
We examine two-sided markets where players arrive stochastically over time and are drawn from a continuum of types. The cost of matching a client and provider varies, so a social planner is faced with two contending objectives: a) to reduce players' waiting time before getting matched; and b) to form efficient pairs in order to reduce matching costs. We show that such markets are characterized by a quick or cheap dilemma: Under a large class of distributional assumptions, there is no `free lunch', i.e., there exists no clearing schedule that is simultaneously optimal along both objectives. We further identify a unique breaking point signifying a stark reduction in matching cost contrasted by an increase in waiting time. Generalizing this model, we identify two regimes: one, where no free lunch exists; the other, where a window of opportunity opens to achieve a free lunch. Remarkably, greedy scheduling is never optimal in this setting.
A system of non-tradable credits that flow between individuals like karma, hence proposed under that name, is a mechanism for repeated resource allocation that comes with attractive efficiency and fairness properties, in theory. In this study, we test karma in an online experiment in which human subjects repeatedly compete for a resource with time-varying and stochastic individual preferences or urgency to acquire the resource. We confirm that karma has significant and sustained welfare benefits even in a population with no prior training. We identify mechanism usage in contexts with sporadic high urgency, more so than with frequent moderate urgency, and implemented as a simple (binary) karma bidding scheme as particularly effective for welfare improvements: relatively larger aggregate efficiency gains are realized that are (almost) Pareto superior. These findings provide guidance for further testing and for future implementation plans of such mechanisms in the real world.
First and foremost, pre-registration is not the all-in-one solution for experimental economics [...]
We define a double auction mechanism, treating in a unified way finite and infinite markets, allowing for ties in reported values, and not imposing any regularity assumptions. It is the first such definition. In all markets, our Double Auction implements market clearing and a Walrasian equilibrium. In finite markets our Double Auction nests as special cases the standard k -Double Auction and in infinite markets the textbook model of continuous and strictly monotone demand and supply. Finally, we establish the convergence of finite to infinite Double Auctions
We conducted a large number of controlled continuous double auction experiments to reproduce and stress-test the phenomenon of convergence to competitive equilibrium under private information with decentralized trading feedback. Our main finding is that across a total of 104 markets (involving over 1,700 subjects), convergence occurs after a handful of trading periods. Initially, however, there is an inherent asymmetry that favors buyers, typically resulting in prices below equilibrium levels. Analysis of over 80,000 observations of individual bids and asks helps identify empirical ingredients contributing to the observed phenomena including higher levels of aggressiveness initially among buyers than sellers.
We provide a game-theoretic analysis of deflationary burn-and-mint tokenomics, studying individual incentives, theoretically and through numerical simulations, to identify the circumstances under which economic collapse can be avoided. We identify a necessary deflation threshold to guarantee preservation of value and propose to fix rewards in fiat as a remedy to the asymptotic underincentivization of network contributors occurring in deflationary burn-and-mint.
How humans behave in repeated strategic interactions, how they learn, how their decisions adapt, and how their decision-making evolves is a topic of fundamental interest in behavioral economics and behavioral game theory [...]