arXiv · 2609.33157
TimelyDAgger: Timing-Aware Expert Querying for VLA Policy Improvement
Abstract
DAgger improves robot policies by aggregating expert supervision from states visited during policy execution. Robot-gated DAgger automates expert queries, allowing the robot to decide when to request expert takeover. While existing gates emphasize detecting the need for assistance, takeover timing also shapes the content of these demonstrations and their value for policy learning. We propose TimelyDAgger, combining Bridge-PCA monitoring of internal vision-language-action (VLA) features with Feedback-guided Threshold Adaptation based on expert behavior to improve takeover timing. We introduce an evaluation framework linking failure detection, takeover timing, and policy improvement, including Target-Aligned Supervision Ratio (TASR) for assessing supervision quality without retraining. Experiments show that takeover timing affects policy learning, with TimelyDAgger achieving competitive failure detection and higher post-training success in most evaluated settings under matched expert-action budgets.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zhixuan Zhao, Peiyan Li, Enhao Zhang, Yueran Tao, Hao Wang, Chenghao Yue, Lei Lv, Wentao Zhao, Jiahao Chen, Xin Liu, Kangyao Huang, Yu Luo, Huaping Liu. 2026-09-27. TimelyDAgger: Timing-Aware Expert Querying for VLA Policy Improvement. https://arxiv.org/abs/2609.33157
Cite the original work for its findings. Save a collection to share your selection of sources.