arXiv · 2610.03333
Equivariant Visual-Tactile Diffusion Policy for Contact-Rich Manipulation
Abstract
Imitation learning for contact-rich manipulation requires high-quality expert data that is expensive to obtain. This makes learning a sample-efficient policy a key issue. To address this, we propose VISTA, a workspace-level equivariant visuotactile diffusion policy for data-efficient contact-rich imitation learning. VISTA projects visual and tactile observations into spherical tokens, injects tactile contact cues into visual spherical directions through permutation-equivariant spherical fusion, and rotates the fused harmonic representation using the end-effector orientation. The resulting representation conditions an equivariant diffusion policy to predict spatially consistent actions. Extensive experiments in both simulation and real-world robotic settings show that VISTA substantially improves data efficiency over strong visuotactile imitation learning baselines. Project website: https://vista-paper.github.io/
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Lik Hang Kenny Wong, Yiyao Ma, Xiu-Shen Wei, Zelong Tan, Zhuheng Song, Dongsheng Xie, Kai Chen, Qi Dou. 2026-10-02. Equivariant Visual-Tactile Diffusion Policy for Contact-Rich Manipulation. https://arxiv.org/abs/2610.03333
Cite the original work for its findings. Save a collection to share your selection of sources.