arXiv · 2609.24563
ARSTAG: An Agentic Real2Sim2Real System for Task-Specific Robot Data Generation
Abstract
Adapting visuomotor policies to new manipulation tasks often requires substantial manual engineering or teleoperated data collection. Simulation can provide task-specific data at scale, but constructing the scene, designing expert behavior, and configuring data generation still require significant per-task effort. We present ARSTAG, an agentic Real2Sim2Real system that turns a single RGB image and a natural-language instruction directly into robot policy-learning data. A hierarchy of language agents constructs a task-scoped simulation scene, generates robot-feasible demonstrations, and expands the training distribution through task-consistent randomization, while a coordinator agent manages cross-stage feedback and recovery. Across seven manipulation tasks spanning grasping, placement, and stacking, the ARSTAG-generated demonstrations enable sim-to-real transfer of three visuomotor policy architectures to a dual-arm robot, with pi0.5 achieving an average real-world success rate of 74.6%. Ablations show that task-consistent randomization substantially improves robustness, and policy performance increases with generated dataset size. Project webpage: https://boweili666.github.io/ARSTAG/.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Bowei Li, Yuner Zhang, Changliu Liu. 2026-09-21. ARSTAG: An Agentic Real2Sim2Real System for Task-Specific Robot Data Generation. https://arxiv.org/abs/2609.24563
Cite the original work for its findings. Save a collection to share your selection of sources.