TacZero: 触覚フィードバックと汎用視覚言語モデルによる訓練不要のペグ挿入
TacZero: Training-Free Peg Insertion Using a General-Purpose Vision-Language Model with Tactile Feedback
汎用視覚言語モデルに触覚情報を入力し、追加学習やタスク固有ルールなしでペグ挿入を実現する手法を提案。実機実験で触覚入力ありが20回中15回成功し、なしの10回を上回った。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Kazutoshi Tanaka
分類: cs.RO
原文アブストラクト
Robots that autonomously determine their actions from language instructions and sensory observations could perform new contact-rich manipulation tasks without task-specific training or hand-designed rules. To perform these tasks, robots must infer how objects contact one another and move as a result, then select actions. For contact inference and action selection, prior approaches involve designing estimation models and tactile feedback control laws, or learning models for object-motion estimation, action-outcome prediction, and action selection from tactile data. Instead, we propose TacZero, which uses a pretrained general-purpose vision-language model (VLM) to interpret visual and tactile observations and select robot actions without additional tactile or manipulation training or task-specific rules for contact interpretation or action selection. TacZero provides the VLM with camera images, robot state, and three-axis tactile responses represented as numerical values or vectors overlaid on the images. From these observations and interaction history, the VLM generates commands specifying target end-effector positions and gripper opening or closing, which a low-level controller executes. In real-world cylindrical-peg insertion experiments, TacZero succeeded in 15 of 20 trials with numerical tactile input, compared with 10 of 20 without tactile input. This study provides a concrete starting point for further research on contact-rich manipulation using general-purpose VLMs and highlights challenges in pursuing this direction.
関連論文
- 密度関数を用いた安全なマルチロボット協調搬送マニピュレーション
- コンパクトなロボットポリシーに必要なのは細粒度の視覚表現マニピュレーション
- ずれた座標系を見抜く:視覚・力覚精密組立のための特権ノイズ蒸留マニピュレーション
- 把持後における物体再配向によるロボット挿入の運動学的修復マニピュレーション
- 未来が一致する間だけコミット:ロボットマニピュレーションのための結果認識型適応アクションチャンキングマニピュレーション
- 油圧クレーンによる丸太山積み処理のための点群からの把持位置学習マニピュレーション