VISTA: 視覚から推定する空間接触注意による高密度接触操作
VISTA: Visually Inferred Spatial ConTact Attention for Contact-Rich Manipulation
接触を伴う精密操作のため、視覚のみからグリッパの変形を推定し、触覚センサなしで接触情報を獲得する模倣学習フレームワークを提案した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Jiayi Chen, Wenlong Dong, Yan Huang, Xianglin Chen, Zijian Lin, Jiaqi Yin, Yushan Liu, Wenbo Ding
分類: cs.RO
原文アブストラクト
Contact-rich manipulation requires precise interaction feedback. While vision-centric imitation learning is prevalent, external visual observations provide indirect and ambiguous cues about contact states, particularly under occlusion or subtle object--gripper interactions; dedicated tactile or force sensors can provide rich contact information but introduce additional hardware complexity, calibration requirements, and deployment costs. To bridge this gap, we propose VISTA-Policy, an imitation learning paradigm that utilizes the Visual Deformation Field (VDF), a 3D displacement representation of a compliant gripper, as high-dimensional visuo-physical feedback. The framework integrates: 1) a Physics-Aware Encoding Engine for real-time VDF decoding; 2) an Energy Aggregation Denoising Mechanism to isolate true interaction signals; and 3) a Deformation-Augmented Policy Network with incremental gripper actions for precise closed-loop correction. Extensive evaluations on Cross-Scale Object Grasping, Cap Unscrewing, and Calligraphy Writing demonstrate that VISTA-Policy outperforms the strong pure-vision baseline 3D Diffusion Policy and the tactile baseline. VISTA-Policy further demonstrates substantial out-of-distribution generalization to unseen object scales and robustness against dynamic disturbances, offering a durable and cost-effective route toward general-purpose fine-grained manipulation in unstructured environments. Project videos and supplementary materials are available at: https://sites.google.com/view/vista-policy.
関連論文
- リー群制約付きMeanFlowによる高速生成的把持マニピュレーション
- PRISM: 投影統合型サンプリングベースMPCとベイズコスト調整による双腕マニピュレーションマニピュレーション
- ロボット操作のための軌道レベル連続行動表現マニピュレーション
- ブラウザ制御と混合ステッパードライバ構成による再現可能な視覚誘導6自由度ロボットアームマニピュレーション
- CSymPlan:高自由度マニピュレータのための認定シンボリックプランニングと制御マニピュレーション
- 力/トルクベースの運動学的適応によるロボット操作タスクマニピュレーション