arXiv Science⌕ Search

arXiv · 2609.37970

PhysWAM: Physically Consistent World Action Model for Autonomous Driving

Abstract

World-action models (WAMs) jointly predict how a scene will evolve and how an agent should act, however joint generation alone does not necessarily impose a shared geometric constraint on these predictions. We present PhysWAM, a unified world-action model for autonomous driving that co-denoises multiview video, metric depth, and ego motion within a single flow-matching transformer. To ground world and action generation in measured scene geometry, we introduce Coupled Point Projection (CPP) that unprojects the generated depth into 3D points, transforms them using the generated $\mathrm{SE}(3)$ ego motion, and minimizes their distance to LiDAR points transformed using the recorded ego motion. This geometric constraint promotes physical consistency with the measured scene by jointly supervising generated depth and motion alongside their standard flow-matching objectives. At inference, trajectory selection relies only on a simple label-free consensus rule, with no learned scorer or simulator feedback. We evaluate PhysWAM across NAVSIM v1 and v2 planning, zero-shot closed-loop transfer, and future video and metric-depth prediction. Despite PhysWAM's simple selection procedure, it achieves strong planning performance and transfers zero-shot to unseen driving environments. It also generates accurate metric depth and temporally coherent video, with CPP improving both planning and depth prediction. Together, these results demonstrate that the geometric relationship between scene depth and ego motion provides a direct way to couple world and action generation within a simple unified model.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Dhruv Parikh, Fengcheng Yu, Quankai Gao, Jiawei Yang, Junjie Ye, Maulik Bhatt, Thang Vu, Charles Ochoa, Rowan McAllister, Igor Vasiljevic, Rajgopal Kannan, Viktor Prasanna, Vitor Guizilini, Yue Wang. 2026-09-29. PhysWAM: Physically Consistent World Action Model for Autonomous Driving. https://arxiv.org/abs/2609.37970

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

An Real-Sim-Real (RSR) Loop Framework for Generalizable Robotic Policy Transfer

The sim-to-real gap remains a critical challenge in robotics, hindering the deployment of algorithms trained in simulation to real-world systems. We propose a flexible Real-to-Sim-to-Real (RSR) framework whose central contribution is an information-theoretic cost function that explicitly accounts for sim-to-real discrepancies. This cost balances two objectives, completing the task and steering the policy to collect real-world samples that are maximally informative for improving transfer. It can be integrated seamlessly into existing reinforcement learning algorithms (e.g., PPO, SAC) and ensures a balanced exploration of critical regions in the real domain. The framework treats differentiable simulation as optional: when a differentiable simulator is available, the collected informative data can also be used to tune simulator parameters. We implement the framework with the MuJoCo MJX platform and demonstrate its generality by evaluating on both manipulation tasks with a 6-DOF robotic arm and locomotion tasks on a legged robot. Empirical results show that our RSR loop yields more efficient data acquisition and substantially improves task performance in real-world that achieves a smoother sim-to-real transfer.

cs.RO↗

Learning-Based Progressive Barrier Control for Robot Manipulators with Initial Errors Outside Prescribed Tracking Bounds

Robot manipulators may start a new task with a tracking error larger than the prescribed tolerance. Conventional barrier controllers generally require the initial error to lie within this tolerance, which prevents their direct use under such conditions. This paper develops a progressive barrier controller that gradually contracts an initial error bound to the required value within a prescribed time. The robot can therefore start outside the final bound, while the direct joint-position error satisfies it after the transition. The closed-form control law combines progressive barrier feedback with an online adaptive torque term based on an actor--critic structure. A Lyapunov analysis establishes bounded closed-loop signals and gives sufficient conditions for satisfaction of the final tracking bound. Two-link simulations consider large initial errors, actuator saturation, dynamic variations, disturbances, and measurement errors. The adaptive term reduces the median root-mean-square tracking error by 45.1\% compared with the zero-weight progressive barrier controller. The simulations also map the initial errors that can be handled at different transition times under fixed torque limits. Finally, an experiment on a Niryo Ned3 Pro illustrates tracking performance using measured position and motor-current data.

cs.RO↗

BIM Informed Visual SLAM for Construction Environments

Monitoring building construction sites requires comparing the as-planned design with the as-built state, which can be estimated in real time using Simultaneous Localization and Mapping (SLAM) techniques. However, visual SLAM is prone to trajectory drift in construction environments, producing maps that are geometrically inaccurate with the actual environment. To address this limitation, we augment an existing RGB-D SLAM system with structural priors derived from the Building Information Model (BIM). The system associates detected walls with their BIM counterparts and includes these correspondences as geometric constraints in the back-end optimization, reducing drift and enhancing global consistency. The proposed method operates in real time and is validated on multiple real construction sites, achieving an average trajectory error reduction of 25.23% and a 7.14% improvement in map accuracy over state-of-the-art baselines. Robustness analyses further demonstrate resilience to incomplete BIM data and geometric discrepancies between as-planned models and the as-built environment.

cs.RO↗