arXiv · 2609.23614
CompVLA: A Variable Compliance Vision-Language-Action Model for Contact-rich Manipulation
Abstract
Contact-rich manipulation, requiring robots to regulate not only motion but also how they yield to external forces, has emerged as the next frontier for Vision-Language-Action (VLA) models. However, existing VLAs output purely kinematic commands, degrading performance on real-world contact-rich tasks. In this paper, we introduce CompVLA, a unified VLA framework that jointly predicts motion and stiffness matrix from RGB and language inputs. Our approach augments the conventional architecture with a dedicated Compliance Expert, which outputs time-varying stiffness and virtual displacement profiles executed via geometric impedance control. We demonstrate that CompVLA achieves the highest average success rate across diverse contact-rich tasks, outperforming both vanilla and compliance-aware VLA baselines, with ablations confirming each component is essential.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jongmin Kim, Junsu Ha, Che-Sang Park, Minchang Song, Hyeokju Jeong, Himchan Hwang, Jianlong Fu, Frank C. Park. 2026-09-20. CompVLA: A Variable Compliance Vision-Language-Action Model for Contact-rich Manipulation. https://arxiv.org/abs/2609.23614
Cite the original work for its findings. Save a collection to share your selection of sources.