arXiv Science⌕ Search

arXiv · 2610.02848

Permutation Robustness Is Not Enough: Action Collapse in Multi-Agent Transformer Policies

Abstract

Transformer policies are attractive for multi-agent robot learning because self-attention can model interactions among agents. However, multi-agent teams are unordered, while transformers typically process agents as ordered token sequences. We study how this mismatch affects cooperative navigation policies under agent-order permutations. Our results show that low permutation error alone can be misleading: policies may appear robust simply because all agents choose the same action. We therefore evaluate policies using both permutation-consistency metrics and action-collapse diagnostics, including action diversity, same-action fraction, and maximum action frequency. A PPO-ID baseline yields non-collapsed behavior but remains order-sensitive, while strong equivariance regularization can still induce homogeneous behavior. A weak equivariance penalty improves the robustness while preserving more diverse actions for teams with \(N=3\) agents, whereas teams with \(N=4\) agents require substantially smaller regularization weights. These findings suggest that multi-agent transformer policies should be evaluated not only by return and permutation robustness, but also by whether they maintain non-collapsed, differentiated multi-agent behavior.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Amit Thakur, Mukesh Singhal. 2026-10-02. Permutation Robustness Is Not Enough: Action Collapse in Multi-Agent Transformer Policies. https://arxiv.org/abs/2610.02848

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Learning Low-Frequency Motion Control for Robust and Dynamic Robot Locomotion

Robotic locomotion is often approached with the goal of maximizing robustness and reactivity by increasing motion control frequency. We challenge this intuitive notion by demonstrating robust and dynamic locomotion with a learned motion controller executing at as low as 8 Hz on a real ANYmal C quadruped. The robot is able to robustly and repeatably achieve a high heading velocity of 1.5 m/s, traverse uneven terrain, and resist unexpected external perturbations. We further present a comparative analysis of deep reinforcement learning (RL) based motion control policies trained and executed at frequencies ranging from 5 Hz to 200 Hz. We show that low-frequency policies are less sensitive to actuation latencies and variations in system dynamics. This is to the extent that a successful sim-to-real transfer can be performed even without any dynamics randomization or actuation modeling. We support this claim through a set of rigorous empirical evaluations. Moreover, to assist reproducibility, we provide the training and deployment code along with an extended analysis at https://articulated.robots.ox.ac.uk/lfmc/.

cs.RO↗

Modelling and Model-Checking a ROS2 Multi-Robot System using Timed Rebeca

Model-based development accelerates prototyping, enables earlier experimentation, and ensures rigorous validation of system design intents. In multi-agent systems with complex asynchronous interactions and concurrency, formal verification, particularly model-checking, offers an automated means of confirming that desired properties hold. Timed Rebeca, an actor-based modelling language supporting reactive, concurrent, and timed behaviors, together with its model-checking tool, provides a powerful framework for this purpose. By leveraging these capabilities, Timed Rebeca can intuitively capture ROS2 node graphs, recurring physical signals, motion primitives, and other time-convertible behaviors. Nevertheless, modelling and verifying multi-robot systems entail significant challenges: abstracting intricate information, bridging the gap between discrete models and continuous system dynamics, and managing large state spaces while preserving fidelity. To address these challenges, we propose discretization strategies tailored to various data types and identify thresholds of abstraction that balance accuracy and tractability. We further introduce optimization techniques to accelerate verification. Our work demonstrates how to systematically design and verify multi-robot systems through Timed Rebeca, efficiently transform continuous dynamics into discrete models for model-checking, and maintain a practical, bidirectional flow between the abstract model and the ROS2 implementation. The accompanying Rebeca and ROS2 codebases, made openly available, serve as a foundational reference for researchers and developers aiming to model and verify advanced autonomous robotic systems.

cs.RO↗

FAME: Force-Adaptive RL for Expanding the Manipulation Envelope of a Full-Scale Humanoid

Maintaining balance under external hand forces is critical for humanoid bimanual manipulation, where interaction forces propagate through the kinematic chain and constrain the feasible manipulation envelope. We propose FAME, a force-adaptive reinforcement learning framework that conditions a standing policy on a learned latent context encoding upper-body joint configuration and bimanual interaction forces jointly, since the base moment a load induces depends on the arm configuration through which it acts. Training applies isotropically sampled 3D forces at each hand under an upper-body pose curriculum, exposing the policy to manipulation-induced perturbations across continuously varying arm configurations. At deployment the interaction force is not measured but reconstructed online from joint torques and states through rigid-body inverse dynamics, requiring no wrist force/torque sensing. We evaluate over $100$ upper-body configurations under swept hand forces, scoring each trial by a task-level criterion that requires the robot both to remain upright and to hold its hands near where the task placed them; all such results run with the estimated force in the loop. At a $150$,mm tolerance FAME reaches $38.9\%$ task success, against $16.6\%$ for a policy given the same force without encoding, $4.3\%$ for a pose-conditioned curriculum policy, and $24.7\%$ for an adversarially trained locomotion policy, which stays upright but recovers by stepping and so relocates the hands. We further demonstrate transfer to task-generated interaction forces in a MuJoCo kitchen environment, and to asymmetric and bimanual loading on a full-scale Unitree H1-2. Code and videos are available on the https://correlllab.github.io/fame_website.

cs.RO↗