arXiv · 2609.33104
VPTwin: Real-Sim-Real Video Prediction for Robotic Manipulation Planning
Abstract
While action-conditioned video prediction provides an intuitive world model for robotics, purely data-driven predictors often suffer from compounding errors and physically implausible hallucinations in long-horizon rollouts, severely undermining downstream action planning. We propose VPTwin, a Real-Sim-Real video prediction framework that anchors real-world future prediction using real-synchronized simulation twins. For a target manipulation task, a VLM reconstructs an executable digital twin from a real demonstration episode. To accommodate the ill-posed estimation of unobserved physical properties, Isaac Sim simulates multiple forward dynamic rollouts across randomized physical configurations under candidate action trajectories. Using these rollouts as in-context references, VPTwin harmonizes both domains, using simulation dynamics to enforce physical plausibility while capturing unmodeled contact interactions from real video. Furthermore, we establish a predictive planning loop using VPTwin to visually verify VLM-proposed actions and guide reliable real-world execution. Evaluations show substantial reductions in physical hallucinations during video prediction and marked improvements in manipulation planning performance.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zhenghao Xiao, Minting Pan, Nantian He, Dongzhan Zhou, Yunbo Wang. 2026-09-27. VPTwin: Real-Sim-Real Video Prediction for Robotic Manipulation Planning. https://arxiv.org/abs/2609.33104
Cite the original work for its findings. Save a collection to share your selection of sources.