arXiv Science⌕ Search

arXiv · 2609.33269

Q-WAM: 4-Bit Quantization of World Action Models with Action-Subspace Protection

Abstract

World Action Models (WAMs) jointly generate video and robot actions through iterative diffusion and perform strongly in robotic manipulation. However, their prohibitive compute and memory costs pose substantial deployment challenges. Post-training quantization (PTQ) can reduce these costs, but existing PTQ methods such as smoothing and rotation are insufficient to maintain the precision of action generation. To overcome this limitation, we propose Q-WAM, a new 4-bit weight-activation quantization for WAMs that preserves the actions the model generates. Specifically, we introduce the \textit{Action Observability Gramian (AOG)}, which measures how much rounding errors in each weighted combination of a layer's input channels change the final action through all denoising steps. We also develop Action-Subspace Protection (ASP), which keeps the few most action-sensitive channel combinations in a tiny 16-bit low-rank branch and quantizes the complementary weights and activations to 4 bits, both as dense matrix multiplications that run efficiently on GPUs. Finally, to preserve action quality with minimal overhead, we identify the experts that matter most for the generated action by aggregating the AOG-derived action mass across the layers of each expert and apply ASP only to those experts. We evaluate Q-WAM on three WAMs, both in simulation and in real-world deployment. On the RoboTwin 2.0 benchmark, it reaches 89.6--93.0\% average success rate, within 1.1 percentage points of the 16-bit models, while reducing the memory of the quantized blocks by 3.1--3.4$\times$. Our method outperforms the strongest baseline, SVDQuant, by 2.5--8.7 percentage points. On a Unitree G1 humanoid and a bimanual UR3 robot, it improves success over SVDQuant by 12.8-17.6 percentage points.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Arash Akbari, Arman Akbari, Jingwu Luo, Yuhao Lei, Yi Gao, Weiwei Chen, Xuan Zhang, Zhenman Fang, Geng Yuan, Yanzhi Wang. 2026-09-27. Q-WAM: 4-Bit Quantization of World Action Models with Action-Subspace Protection. https://arxiv.org/abs/2609.33269

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Fading Expert Guidance: Bridging Model-Based and Learning-Based Control for Abortable Autonomous Overtaking

Overtaking on two-lane roads is a safety-critical decision-making problem for autonomous vehicles, since oncoming traffic may force the ego vehicle to abort and merge back to its original lane. Deep reinforcement learning (DRL) is promising for such continuous-control tasks, but it requires substantial interaction data and may produce unsafe exploratory actions during training. Model-based controllers provide structured behavior, but depend on model fidelity and hand-designed logic. This paper proposes a fading expert-guidance framework that uses a model-based controller to guide DRL training for autonomous overtaking. The expert combines a constrained iterative LQR (CiLQR) planner with PID-based auxiliary controllers, activated when the optimized plan becomes infeasible. The expert action enters the actor objective through a fading term, interpreted as a time-varying soft trust region around the expert policy. This term first biases the policy toward the expert and then vanishes, allowing optimization according to the reinforcement-learning objective. The method is evaluated with PPO, TD3, and SAC. Simulations with bidirectional traffic show improved sample efficiency and final task performance. The safety effects are algorithm-dependent: guidance reduces vehicle collisions for SAC, eliminates boundary collisions for TD3, and leaves the PPO collision profile unchanged. Overall, the results support fading expert guidance as a practical mechanism for transferring model-based driving knowledge to learning-based controllers without restricting the final policy to imitation in an abortable overtaking task with bidirectional traffic.

cs.RO↗

BR-MPPI: Barrier-Rate Guided MPPI for Enforcing Multiple Inequality Constraints with Learned Signed Distance Fields

Model Predictive Path Integral (MPPI) control provides a sampling-based framework for optimal control, while Control Barrier Functions (CBFs) provide a principled means of enforcing safety constraints. We introduce BR-MPPI, which integrates CBF-like conditions into MPPI's control sampling procedure. CBFs impose inequality constraints that bound the rate of change of barrier functions using a class-K function of the barrier value. We instead impose the CBF condition as an equality constraint using a parametric linear class-K function and augment the system state with its parameter. The parameter's time derivative serves as an additional control input optimized by MPPI. We further design a cost function that promotes parameter values consistent with Nagumo's condition at the safe-set boundary, thereby encouraging safety. The resulting multiple state- and control-dependent equality constraints pose a challenge for random control sampling. We address this through state transformations and control projections inspired by manifold path planning that map sampled controls onto the constraint manifold. We also incorporate learned signed distance fields to represent robot geometry and reduce computation time. Simulations demonstrate improved sample efficiency over vanilla MPPI and higher navigation success rates across five robot models compared with safety-oriented MPPI variants. Hardware experiments on a quadrotor further demonstrate the method's ability to navigate constrained environments near safe-set boundaries.

cs.RO↗

Tunable Leg Stiffness in a Monopedal Hopper for Energy-Efficient Vertical Hopping Across Varying Ground Profiles

We present the design and implementation of HASTA (Hopper with Adjustable Stiffness for Terrain Adaptation), a vertical hopping robot with real-time tunable leg stiffness, aimed at optimizing energy efficiency across various ground profiles (a pair of ground stiffness and damping conditions). By adjusting leg stiffness, we aim to maximize apex hopping height, a key metric for energy-efficient vertical hopping. We hypothesize that softer legs perform better on soft, damped ground by minimizing penetration and energy loss, while stiffer legs excel on hard, less damped ground by reducing limb deformation and energy dissipation. Through experimental tests and simulations, we find the best leg stiffness within our selection for each combination of ground stiffness and damping, enabling the robot to achieve maximum steady-state hopping height with a constant energy input. These results support our hypothesis that tunable stiffness improves energy-efficient locomotion in controlled experimental conditions. In addition, the simulation provides insights that could aid in the future development of controllers for selecting leg stiffness.

cs.RO↗