arXiv · 2610.09514
STRIKE: Learning Visual State Transitions for Physical World Modeling
Abstract
Physical world modeling requires predicting how interactions change a scene, not merely generating coherent motion. We propose STRIKE, a framework that separates visual state transition learning from dense video generation. We construct event-aligned supervision by extracting observed states from training videos and pairing them with transition descriptions and temporal offsets. An image-based transition model learns to predict the next scene configuration from the current image, a local transition specification, and elapsed time. At inference, a pretrained vision-language planner predicts time transition specifications, and recursive application of the learned transition model produces a sequence of future visual states. A separately trained dynamic model then generates the complete rollout conditioned on these states and their temporal locations. Experiments on Physics-IQ Verified, PhyGenBench, Pisa-Experiments, and RoboTwin2.0 show improvements of STRIKE over the corresponding video-backbone baselines in benchmark measures of physical consistency and manipulation-video fidelity. These results support learned visual state transitions as an effective intermediate representation for physical world modeling.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wenbin Teng, Tianshuo Xu, Depu Meng, Yuelei Li, Quentin Herau, Yihan Hu, Yajie Zhao, Wei Zhan. 2026-10-07. STRIKE: Learning Visual State Transitions for Physical World Modeling. https://arxiv.org/abs/2610.09514
Cite the original work for its findings. Save a collection to share your selection of sources.