arXiv · 2609.12898
UniPart: Towards Zero-shot Language-Grounded 3D Part Segmentation for Embodied Interaction
Abstract
Fine-grained robotic manipulation depends on understanding parts, not only whole objects. Existing 3D foundation models tend to be either generalized but object-aware, or part-aware but limited to closed-set taxonomies, which weakens zero-shot transfer. We study text-conditioned 3D part segmentation, where a free-form phrase selects a functional part on point cloud. We introduce UniPart, a feed-forward cross-modal 3D Transformer that conditions CLIP text embedding. To scale supervision, we build LangPart-1M with 160K+ Objaverse assets and 8M text to part pairs using multi-view consistent part generation. We further manually label a high-quality subset, LangPart-4K, for fine-tuning and evaluation. UniPart achieves strong zero-shot results on open-vocabulary part benchmarks and transfers to language-conditioned part grasping in real world.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xinqiang Yu, Zekun qi, Jiawei He, Wenyao Zhang, Xuchuan Chen, Guaocai Yao, Li Yi, Zhaoxiang Zhang, He Wang. 2026-09-11. UniPart: Towards Zero-shot Language-Grounded 3D Part Segmentation for Embodied Interaction. https://arxiv.org/abs/2609.12898
Cite the original work for its findings. Save a collection to share your selection of sources.