arXiv · 2609.32155
RecastVLA: From Past Interaction to Future Control with Adaptive Policy States
Abstract
Sequential manipulation requires a robot to track what has already happened, even when the current scene no longer reveals it. Policies with explicit history representations make past interactions available as context for current decisions. We ask how action generation itself can form a persistent state for subsequent control. Building on action-side test-time training, RecastVLA maintains an adaptive policy state within a flow-matching vision-language-action policy. The state is represented by shared fast weights and remains fixed throughout action generation. Depth-specific interfaces read the same state, while features across depths and flow evaluations jointly define one update for the next policy call. Subsequent action losses train the initialization, interfaces, and update rule by differentiating through earlier state transitions. At deployment, updates use the policy's own action-generation features without expert action labels. Across LIBERO, RoboTwin, RoboDojo, and twelve real-robot tasks, RecastVLA improves mean success over a matched policy trained without test-time training, including 10.68 percentage points on RoboTwin Clean-to-Clean. In controlled RoboTwin comparisons, retaining state improves success, and the shared design exceeds independently trained layer-local TTT by 2.58 points.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wenbo Li, Jun Yang, Yiteng Chen, Wei Zhang, Qingyao Wu. 2026-09-26. RecastVLA: From Past Interaction to Future Control with Adaptive Policy States. https://arxiv.org/abs/2609.32155
Cite the original work for its findings. Save a collection to share your selection of sources.