IEEE INFOCOM 2025-IEEE CONFERENCE ON COMPUTER COMMUNICATIONS(2025)
Univ Victoria
被引用0|浏览7
摘要
In this paper, we study unknown-game bandits, where multiple agents play a general-sum game repeated over T rounds. In each round, each agent independently selects an action and observes the reward for that action. The game is unknown to every agent, meaning each agent has no knowledge about the underlying game structure, the number of other agents, or their actions and rewards. Such unknown-game bandits have wide applications in computer and communication networks, including congestion control and network selection. The goal of each agent is to minimize swap regret, which measures the performance gap from a broader class of competitors than the traditional external regret that only compares against competitors always playing a fixed action. Our main contribution is to bridge the gap in the literature by proving the first swap-regret bound with a time-dependence of (O) over tilde (T-1/4) if the proposed learning algorithm based on optimistic follow-the-regularized-leader (OFTRL) is played by all agents involved in the game, where (O) over tilde(center dot) hides log-arithmic factors. This regret bound demonstrates a faster convergence rate with respect to the number of rounds T compared to the state-of-the-art swap regret bound of O(T-1/2). Furthermore, we demonstrate the efficacy of the proposed algorithm through an application in heterogeneous network selection with both numerical and simulation-based experiments.
更多
查看译文
关键词
Faster Convergence,Learning Algorithms,Gap In The Literature,Convergence Rate,Communication Network,Numerical Experiments,Number Of Agents,Heterogeneous Network,Multiple Agents,Goal Of The Agent,Heterogeneous Selection,Time And Space,Throughput,Wireless Networks,Online Learning,Joint Action,Space Complexity,Mixed Strategy,Equilibrium Of The Game,Multi-armed Bandit,Bandit Problem,Stronger Notion