Intelligent metasurfaces are often modeled as near-perfect passive beamformers. In practice, their gains require channel information that is hard to obtain and maintain. This paper provides a tractable analysis of metasurface gains in cellular networks under imperfect control. We first study single-surface power scaling under realistic conditions, including bounded alignment errors and spatially correlated fading. Near-coherent operation can yield a quadratic gain with the number of surface reflecting elements, while noncoherent operation yields a linear gain. Embedding these behaviors in a stochastic-geometry network model reveals two operating regimes: a beamforming-dominant regime, where the reflected path can be strong and stable, and a diffusing-dominant regime, where reflections aggregate almost randomly and the direct link prevails. We capture the transition via a moment-matched Gamma model for the useful-signal power, derive downlink SINR coverage, and quantify the roles of metasurface density, surface size, and coherence quality. Finally, mapping coverage to TCP CUBIC throughput identifies when physical-layer improvements might translate into end-to-end gains.
The paper focuses on a particular polling system known as the cyclic Bernoulli polling (CBP) system, where a server moves cyclically between the stations and serves the queue at a station with a certain probability when polled. Each station follows either a gated or partially exhaustive service discipline. In the steady state of such a system, we study a new game-theoretic aspect, where, the stations strategically choose the probability of accepting or rejecting the service from the server when polled. We examine three variants of non-cooperative games among stations: (i) each station selfishly minimizes its expected waiting time, (ii) a team game where each station minimizes the expected workload of the system, and (iii) stations act with partial cooperation, incurring an additional linear cost. We begin by presenting a new result for the CBP system regarding the continuity of expected waiting times in relation to the probabilities selected by the stations. For each game, we then investigate the existence and uniqueness of the Nash equilibrium (NE). In some cases, the NE is explicitly derived, while in others, characterizing the NE remains challenging due to the complex dependence of waiting times on the non-trivial buffer occupancy equations. Nonetheless, we analyze the NE and its properties through numerical experiments. Notably, in many instances, stations opt to accept service with a probability less than 1—a trend observed even among selfish stations.
The Kelly or proportional allocation mechanism is a simple and efficient auction-based scheme that distributes an infinitely divisible resource proportionally to the agents bids. When agents are aware of the allocation rule, their interactions form a game extensively studied in the literature. This paper examines the less explored repeated Kelly game, focusing mainly on utilities that are logarithmic in the allocated resource fraction. We first derive this logarithmic form from fairness-throughput trade-offs in wireless network slicing, and then prove that the induced stage game admits a unique Nash equilibrium NE. For the repeated play, we prove convergence to this NE under three behavioral models: (i) all agents use Online Gradient Descent (OGD), (ii) all agents use Dual Averaging with a quadratic regularizer (DAQ) (a variant of the Follow-the-Regularized leader algorithm), and (iii) all agents play myopic best responses (BR). Our convergence results hold even when agents use personalized learning rates in OGD and DAQ (e.g., tuned to optimize individual regret bounds), and they extend to a broader class of utilities that meet a certain sufficient condition. Finally, we complement our theoretical results with extensive simulations of the repeated Kelly game under several behavioral models, comparing them in terms of convergence speed to the NE, and per-agent time-average utility. The results suggest that BR achieves the fastest convergence and the highest time-average utility, and that convergence to the stage-game NE may fail under heterogeneous update rules.
In the multi-resource Kelly mechanism, players obtain a share of each resource in proportion to the bids they place on it. They thus engage in a non-cooperative game where they distribute their budgets across multiple resources. In this paper, we study the repeated variant of this game under standard no-regret algorithms, namely Online Gradient Descent (OGD) and Dual Averaging (DA) algorithms. More specifically, we investigate an additive utility framework with heterogeneous valuations across resources, where each resource-specific utility can be either logarithmic or linear. In this setting, we prove uniqueness of the Nash equilibrium. Moreover, we prove convergence of OGD and DA in the repeated game to this unique Nash Equilibrium. Extensive numerical simulations validate the theoretical results and measure convergence speed across different settings.
Join-the-shortest queue (JSQ) and its variants have often been used in solving load balancing problems. The aim of such policies is to minimize the average system occupation, e.g., the customer's system time. In this paper, we extend the load balancing setting to include constraints that may be imposed, e.g., due to the communication network. First, we cast the problem in the framework of constrained MDPs: this permits us to address both action-dependent constraints, such as, e.g, bandwidth limitation, and state-dependent constraints, such as, e.g., minimum queue utilization. Hence, unlike the state-of-the-art approaches in load balancing, we derive new policies that satisfy the constraints while minimizing system occupancy. Extensive numerical simulations have evaluated their performance under various system settings.
We consider a peer-to-peer electricity market modeled as a private network game, where end users minimize their cost by computing their demand and controllable generation. Their nominal demand constitutes sensitive information that they might want to keep private. We prove that the private network game admits a unique variational equilibrium, which depends on the private information of all end users. Thus, to update their strategy, end users rely on randomized readings. A data aggregator is introduced, which aims to learn the end users' private information, while remunerating them depending on the quality of their readings. Using performative prediction, we define a decision-dependent game explicitly taking into account the distribution shift caused by the end users' hidden ability. The decision-dependent game coincides with a Stackelberg game when the end users' hidden abilities are best responses. Further, the market robustness can be quantified by evaluating the efficiency loss as the difference between the social cost in the performatively stable equilibrium and the optimum. We show that under mild assumptions, the performatively stable equilibrium can be found by distributed and sequential variants of the repeated stochastic gradient method while we propose a two-timescale stochastic approximation method to learn Stackelberg equilibrium. Finally, we formulate the data aggregator's optimal contract design as a bilevel optimization problem that we cast as a more tractable nonlinear nonconvex optimization problem which can be solved using simulated annealing. Simulations on small and large scale problem instances illustrate the results.
In edge computing systems, autonomous agents must make fast local decisions while competing for shared resources. Existing MARL methods often resume to centralized critics or frequent communication, which fail under limited observability and communication constraints. We propose a decentralized framework in which each agent solves a constrained Markov decision process (CMDP), coordinating implicitly through a shared constraint vector. For the specific case of offloading, e.g., constraints prevent overloading shared server resources. Coordination constraints are updated infrequently and act as a lightweight coordination mechanism. They enable agents to align with global resource usage objectives but require little direct communication. Using safe reinforcement learning, agents learn policies that meet both local and global goals. We establish theoretical guarantees under mild assumptions and validate our approach experimentally, showing improved performance over centralized and independent baselines, especially in large-scale settings.
Consider an M/M/1-type queue where joining attains a known reward, but a known waiting cost is paid per time unit spent queueing. In the 1960 s, Naor showed that any arrival optimally joins the queue if its length is less than a known threshold. Yet acquiring knowledge of the queue length often brings an additional cost, e.g., website loading time or data roaming charge. Therefore, our model presents any arrival with three options: join blindly, balk blindly, or pay a known inspection cost to make the optimal joining decision by comparing the queue length to Naor’s threshold. In a recent paper, Hassin and Roet-Green prove that a unique Nash equilibrium always exists and classify regions where the equilibrium probabilities are nonzero. We complement these findings with new closed-form expressions for the equilibrium probabilities in the majority of cases. Further, Hassin and Roet-Green show that minimising inspection cost maximises social welfare. Envisaging a queue operator choosing where to invest, we compare the effects of lowering inspection cost and increasing the queue-joining reward on social welfare. We prove that the former dominates and that the latter can even have a detrimental effect on social welfare.
The Kelly mechanism is a proportional allocation auction widely adopted in decentralized resource allocation systems to share an infinitely divisible resource among competing agents. We analyze the sequential game it induces when agents have α -fair utilities and behave strategically. Our main result proves that synchronous best-response updates drive bids to the unique Nash equilibrium at a linear rate for α∈{0,1,2} . Extensive simulations reveal that best-response dynamics reach equilibrium significantly faster than previously proposed no-regret learning algorithms.
The Kelly or proportional allocation mechanism is a simple and efficient auction-based decentralized resource allocation scheme that distributes an infinitely divisible resource proportionally to the agents' bids. When agents are aware of the allocation mechanism, their interactions form a game. The properties of its Nash equilibria are well understood under the simplifying assumption of unbounded budgets. In this paper, we analyze the game in a more realistic budget-constrained setting, motivated by its optimality in terms of the liquid price of anarchy (LPoA). Specifically, we establish a sufficient condition for the uniqueness of the Nash equilibrium and design a distributed sequential learning procedure that provably converges to the equilibrium. In particular, our sufficient condition holds when the payoff functions of the agents are of the proportional fair type in the allocated fraction. Finally, extensive numerical experiments shed light on the interplay between the heterogeneity of the payoff functions and the agents' budgets.
This paper introduces constrained correlated equilibrium, a solution concept combining correlation and coupled constraints in finite non-cooperative games. In the general case of an arbitrary correlation device and coupled constraints in the extended game, we study the conditions for equilibrium. In the particular case of constraints induced by a feasible set of probability distributions over action profiles, we first show that canonical correlation devices are sufficient to characterize the set of constrained correlated equilibrium distributions and provide conditions of their existence. Second, it is shown that constrained correlated equilibria of the mixed extension of the game do not lead to additional equilibrium distributions. Third, we show that the constrained correlated equilibrium distributions may not belong to the polytope of correlated equilibrium distributions. Finally, we illustrate these results through numerical examples.
The worst-case data-generating (WCDG) probability measure is introduced as a tool for characterizing the generalization capabilities of machine learning algorithms. Such a WCDG probability measure is shown to be the unique solution to two different optimization problems: (a) The maximization of the expected loss over the set of probability measures on the datasets whose relative entropy with respect to a reference measure is not larger than a given threshold; and (b) The maximization of the expected loss with regularization by relative entropy with respect to the reference measure. Such a reference measure can be interpreted as a prior on the datasets. The WCDG cumulants are finite and bounded in terms of the cumulants of the reference measure. To analyze the concentration of the expected empirical risk induced by the WCDG probability measure, the notion of (epsilon, delta)-robustness of models is introduced. Closed-form expressions are presented for the sensitivity of the expected loss for a fixed model. These results lead to a novel expression for the generalization error of arbitrary machine learning algorithms. This exact expression is provided in terms of the WCDG probability measure and leads to an upper bound that is equal to the sum of the mutual information and the lautum information between the models and the datasets, up to a constant factor. This upper bound is achieved by a Gibbs algorithm. This finding reveals that an exploration into the generalization error of the Gibbs algorithm facilitates the derivation of overarching insights applicable to any machine learning algorithm.
Modern portable devices can execute increasingly sophisticated AI models on sensed data. The complexity of such processing tasks is data-dependent and has relevant energy cost. This work develops an Age of Information markovian model for a system where multiple battery-operated devices perform data processing and energy harvesting in parallel. Part of their computational burden is offloaded to an edge server which polls devices at given rate. The structural properties of an optimal policy for a single device-server system are derived. They permit to define a new model-free reinforcement learning method specialized for monotone policies, namely Ordered Q-Learning, providing a fast procedure to learn the optimal policy. The method is oblivious to the devices' battery capacities, the cost and the value of data batch processing and to the dynamics of the energy harvesting process. Finally, the polling strategy of the server is optimized by combining this policy improvement technique with stochastic approximation methods. Extensive numerical results provide insight into the system properties and demonstrate that the proposed learning algorithms outperform existing baselines.
In this paper, the worst-case probability measure over the data is introduced as a tool for characterizing the generalization capabilities of machine learning algorithms. More specifically, the worst-case probability measure is a Gibbs probability measure and the unique solution to the maximization of the expected loss under a relative entropy constraint with respect to a reference probability measure. Fundamental generalization metrics, such as the sensitivity of the expected loss, the sensitivity of the empirical risk, and the generalization gap are shown to have closed-form expressions involving the worst-case data-generating probability measure. Existing results for the Gibbs algorithm, such as characterizing the generalization gap as a sum of mutual information and lautum information, up to a constant factor, are recovered. A novel parallel is established between the worst-case data-generating probability measure and the Gibbs algorithm. Specifically, the Gibbs probability measure is identified as a fundamental commonality of the model space and the data space for machine learning algorithms.
The deployment of RISs in future 6G networks is expected to substantially improve mobile network coverage. This paper introduces a new cross-layer low-complexity scheme for the online optimization of RIS-assisted communication systems. It jointly combines BS and RIS configuration and fair UEs’ scheduling which is critical for high-performance deployments. A RIS beam synthesis method is especially proposed for RIS configuration. The proposed solution embeds two nested control loops: i) a fast control loop working at the OFDMA slot scale and consisting in a standard UEs proportional fair scheduler, and ii) a slow control loop operating at the OFDMA frame scale which adapts the RIS’ configuration to the UEs’ spatial distribution and maximizes the UEs’ aggregated performance. The slow control loop is based on an online stochastic approximation algorithm whose convergence to the optimal restpoint is proved. In a reference scenario, the proposed scheduler achieves a gain of 47% in mean spectral efficiency for NLOS UEs over a baseline scheme.
We model in this paper the multipopulation vaccinatinon game over a fully connected graph. Each player decides whether to purchase a vaccine or not, and if they do, then they further decide which vaccine to purchase among a finite number of vaccine producers. The players need not be indistinguishable. A potential consumer belongs to a risk type that characterizes how important it is for them to be vaccinated. The cost of a vaccine may depend on the demand, on the cost of the production, and on the consumer’s class. We prove in the existence of an equilibrium within pure policies in the general multipopulation case. We further derive some properties of the equilibria in the case of a single risk-class.
We consider the inherent timeline structure of the appearance of content in online social networks (OSNs) while studying content propagation. We model the propagation of a post/content of interest by an appropriate multi-type branching process. The branching process allows one to predict the emergence of global macro properties (e.g., the spread of a post in the network) from the laws and parameters that determine local interactions. The local interactions largely depend upon the timeline (an inverse stack capable of holding many posts and one dedicated to each user) structure and the number of friends (i.e., connections) of users, etc. We explore the use of multi-type branching processes to analyze the viral properties of the post, e.g., to derive the expected number of shares, the probability of virality of the content, etc. In OSNs, the new posts push down the existing contents in timelines, which can greatly influence content propagation; our analysis considers this influence. We find that one leads to draw incorrect conclusions when the timeline (TL) structure is ignored: (a) for instance, even less attractive posts are shown to get viral; (b) ignoring TL structure also indicates erroneous growth rates. More importantly, one cannot capture some interesting paradigm shifts/phase transitions; for example, virality chances are not monotone with network activity parameter, as shown by analysis including TL influence. In the last part, we integrate the online auctions into our viral marketing model. We study the optimization problem considering real-time bidding. We again compared the study with and without considering the TL structure for varying activity levels of the network. We find that the analysis without TL structure fails to capture the relevant phase transitions, thereby making the study incomplete.
We consider a distributed computing setting wherein a central entity seeks power from computational providers by offering a certain reward in return. The computational providers are classified into long-term stakeholders that invest a constant amount of power over time and players that can strategize on their computational investment. In this paper, we model and analyze a stochastic game in such a distributed computing setting, wherein players arrive and depart over time. While our model is formulated with a focus on volunteer computing, it equally applies to certain other distributed computing applications such as mining in blockchain. We prove that, in Markov perfect equilibrium, only players with cost parameters in a relatively low range which collectively satisfy a certain constraint in a given state, invest. We infer that players need not have knowledge about the system state and other players' parameters, if the total power that is being received by the central entity is communicated to the players as part of the system's protocol. If players are homogeneous and the system consists of a reasonably large number of players, we observe that the total power received by the central entity is proportional to the offered reward and does not vary significantly despite the players' arrivals and departures, thus resulting in a robust and reliable system. We then study by way of simulations and mean field approximation, how the players' utilities are influenced by their arrival and departure rates as well as the system parameters such as the reward's amount and dispensing rate. We observe that the players' expected utilities are maximized when their arrival and departure rates are such that the average number of players present in the system is typically between 1 and 2, since this leads to the system being in the condition of least competition with high probability. Further, their expected utilities increase almost linearly with the offered reward and converge to a constant value with respect to its dispensing rate. We conclude by studying a Stackelberg game, where the central entity decides the amount of reward to offer, and the computational providers decide how much power to invest based on the offered reward.
This paper considers different pricing models for a platform based rental system, such as Airbnb. A linear model is assumed for the demand response to price, and existence and uniqueness conditions for Nash equilibria are obtained. The Stackelberg equilibrium prices for the game are also obtained, and an iterative scheme is provided, which converges to the Nash equilibrium. Different cooperative pricing schemes are studied, and splitting of revenues based on the Shapley value is discussed. It is shown that a division of revenue based on the Shapley value gives a revenue to the platform proportional to its control of the market. The demand response function is modified to include user response to quality of service. It is shown that when the cost to provide quality of service is low, both renter and the platform will agree to maximize the quality of service. However, if this cost is high, they may not always be able to agree on what quality of service to provide.
Sara Alouf合作论文数INRIA Sophia Antipolis - Projet MAESTRO;2004 Route des Lucioles28
Bruno Gaujal合作论文数LIG;INRIA 27
Hisao Kameda (亀田壽夫)合作论文数University of Tsukuba19
Dieter Fiems合作论文数SMACS Research Group
Department of telecommunications and information processing (TW07)10