arXiv · 2609.33595
Beyond One-Step Accuracy: State-Affine Latent Transition for Reliable Visual Planning
Abstract
Joint-embedding world models enable visual planning by learning action-conditioned dynamics in latent space. Yet they are commonly trained for one-step prediction on encoded states, while planning recursively applies the learned transition to its own predictions. One-step accuracy therefore does not capture how prediction errors propagate under recursive rollout. We decompose multi-step rollout error into the errors introduced at individual steps and their propagation through subsequent transitions. We show that state-affine dynamics are precisely the differentiable transitions with state-independent Jacobians, eliminating the nonlinear propagation residual and making the error propagation operators depend only on the action sequence. Guided by this result, we introduce SALT (State-Affine Latent Transition), an action-conditioned state-affine dynamics model in which the action modulates both the state transformation and the additive update. We train SALT through recursive multi-step rollout supervision, feeding each predicted latent state back into the transition so that training matches how the model is used during planning. Across four visual planning environments, SALT exhibits $1.48$--$2.19\times$ higher one-step prediction error than the matched LeWM baseline, yet improves closed-loop success in every environment by $10.0$ percentage points on average. On OGBench-Cube, the fraction of episodes that fail with a sharp rise in model-predicted cost after execution decreases from $23.3%$ to $2.0%$.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Boyuan Zhang, Yingjun Du, Xiantong Zhen, Ling Shao. 2026-09-27. Beyond One-Step Accuracy: State-Affine Latent Transition for Reliable Visual Planning. https://arxiv.org/abs/2609.33595
Cite the original work for its findings. Save a collection to share your selection of sources.