Machine Learning and Knowledge Discovery in Databases Research Track(2026)
Université Paris-Saclay
被引用0|浏览0
摘要
The Centralized Training with Decentralized Execution (CTDE) paradigm has become increasingly popular in multi-agent reinforcement learning and is widely adopted in recent works. However, decentralized policies operate under partial observations and may achieve suboptimal performance compared to centralized policies, while naive centralized policies often struggle to scale with larger numbers of agents. To address these limitations, we introduce Centralized Permutation Equivariant (CPE) learning, a scalable centralized training and execution (CTE) framework that transforms standard CTDE algorithms by replacing decentralized execution with a fully centralized policy, while retaining their centralized training components. Our policy network is built upon a Global–Local Permutation Equivariant (GLPE) architecture, which is lightweight, computationally efficient, and agent-number-agnostic at the architectural level. Empirical results show that CPE can be seamlessly integrated with both value decomposition and actor–critic methods, consistently improving the performance of classical CTDE approaches across cooperative benchmarks such as MPE, SMAC, and RWARE, while matching state-of-the-art performance on RWARE. These results suggest that scalable centralized execution constitutes a promising alternative to decentralized policies in coordination-intensive MARL settings.