CompVLA: 接触を伴うマニピュレーションのための可変コンプライアンス視覚-言語-行動モデル
CompVLA: A Variable Compliance Vision-Language-Action Model for Contact-rich Manipulation
RGB画像と言語入力から動作と剛性行列を同時に予測し、接触の多いタスクで高い成功率を達成するVLAフレームワークを提案した。
著者: Jongmin Kim, Junsu Ha, Che-Sang Park, Minchang Song, Hyeokju Jeong, Himchan Hwang, Jianlong Fu, Frank C. Park
分類: cs.RO
原文アブストラクト
Contact-rich manipulation, requiring robots to regulate not only motion but also how they yield to external forces, has emerged as the next frontier for Vision-Language-Action (VLA) models. However, existing VLAs output purely kinematic commands, degrading performance on real-world contact-rich tasks. In this paper, we introduce CompVLA, a unified VLA framework that jointly predicts motion and stiffness matrix from RGB and language inputs. Our approach augments the conventional architecture with a dedicated Compliance Expert, which outputs time-varying stiffness and virtual displacement profiles executed via geometric impedance control. We demonstrate that CompVLA achieves the highest average success rate across diverse contact-rich tasks, outperforming both vanilla and compliance-aware VLA baselines, with ablations confirming each component is essential.