MAGRPO: Accelerated MARL Training for Fluid Antenna-Assisted Wireless Network Optimization
Fluid antenna systems (FASs) improve wireless links by repositioning antenna elements to exploit favorable spatial channel variations. Jointly optimizing fluid antenna (FA) positions, beamforming, and transmit power in a multi-cell network is challenging because the problem is non-convex and each base station has only local information during decentralized execution. However, the representative multi-agent reinforcement learning (MARL) algorithms, namely the on-policy multi-agent proximal policy optimization (MAPPO) and the off-policy multi-agent twin delayed deep deterministic policy gradient (MATD3), suffer from excessively long training times. To address this challenge, we formulate the problem as a decentralized partially observable Markov decision process (Dec-POMDP) and propose multi-agent group relative policy optimization (MAGRPO) under centralized training with decentralized execution. MAGRPO constructs relative advantages from groups of joint trajectories, thereby eliminating the centralized critic and generalized advantage estimation used by MAPPO; under parameter sharing, this critic-free design reduces the per-step computational complexity by approximately half. Simulations show that joint FA optimization provides several-fold sum-rate gains over fixed-position configurations. Across the evaluated antenna settings, MAGRPO achieves test sum rates higher than those of MATD3 and comparable to or slightly higher than those of MAPPO. Meanwhile, it reduces the training time by 20%-23% compared with MAPPO and by about 35% compared with MATD3.