arXiv · 2609.14268
Learning Communication-Conditioned Generative Policies for Decentralized Multi-Agent Collision Avoidance
Abstract
In this work, we propose a decentralized communication-conditioned generative framework for multi-agent collision avoidance. Agents generate short-horizon action sequences using a flow-matching policy trained from privileged offline demonstrations with access to global state. The demonstrations do not include explicit communication signals; instead, agents learn to exchange and aggregate latent messages that encode interaction-relevant intent under partial observability. This formulation supports flexible inference at test time, where unconditioned generation corresponds to independent behavior and communication-conditioned generation enables coordinated interaction without centralized planning. The resulting policies operate in a fully decentralized manner at execution time, relying only on local observations and learned messages. Combined with a receding-horizon inference scheme, the proposed approach enables efficient single-step inference of short-horizon action sequences and degrades gracefully under communication dropouts. Extensive simulation results demonstrate near-expert collision avoidance performance and strong generalization to denser, unseen multi-agent scenarios, along with zero-shot transfer to real-robot experiments.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Prajwal Koirala, Mark Campbell. 2026-09-13. Learning Communication-Conditioned Generative Policies for Decentralized Multi-Agent Collision Avoidance. https://arxiv.org/abs/2609.14268
Cite the original work for its findings. Save a collection to share your selection of sources.