日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.07621

TacZero: 触覚フィードバックと汎用視覚言語モデルによる訓練不要のペグ挿入

TacZero: Training-Free Peg Insertion Using a General-Purpose Vision-Language Model with Tactile Feedback

シェア:XThreadsFacebookLINEはてブBluesky

汎用視覚言語モデルに触覚情報を入力し、追加学習やタスク固有ルールなしでペグ挿入を実現する手法を提案。実機実験で触覚入力ありが20回中15回成功し、なしの10回を上回った。

詳しい要約

1. どんなもの?

- 言語指示と視覚・触覚観測からロボットが自律的に行動を決定する、訓練不要の contact-rich manipulation 手法 TacZero を提案。 - 事前学習済み general-purpose vision-language model (VLM) を用い、追加の触覚・操作訓練やタスク固有ルールなしで peg insertion を実行。 - カメラ画像、ロボット状態、3軸 tactile responses を数値または画像上に重ねたベクトルとして VLM に与える。 - VLM が target end-effector positions と gripper の開閉コマンドを生成し、low-level controller が実行する。

2. 先行研究と比べてどこがすごい?

- 従来は contact inference と action selection のために estimation models や tactile feedback control laws を設計、または tactile data から object-motion estimation・action-outcome prediction・action selection を学習していた。 - TacZero は追加の触覚・操作訓練やタスク固有ルールを必要とせず、pretrained general-purpose VLM をそのまま利用する点が異なる。 - 実世界の cylindrical-peg insertion で、触覚入力ありは20回中15回成功、触覚入力なしは20回中10回成功。

3. 技術・手法の肝は?

- pretrained general-purpose vision-language model (VLM) に camera images、robot state、three-axis tactile responses を数値または画像上に重ねたベクトルとして入力。 - これらの観測と interaction history から、VLM が target end-effector positions と gripper opening/closing を指定するコマンドを生成。 - 生成されたコマンドを low-level controller が実行する。 - 触覚・操作に関する追加訓練やタスク固有の contact interpretation/action selection ルールは用いない。

4. どうやって有効だと検証した?

- 実世界の cylindrical-peg insertion 実験を実施。 - 数値 tactile input ありの場合、20試行中15回成功。 - tactile input なしの場合、20試行中10回成功。 - この比較により、触覚入力の有無が成功率に影響することが示された。

5. 議論はある?

- 本研究は general-purpose VLMs を用いた contact-rich manipulation のさらなる研究の具体的な出発点を提供する。 - この方向性を追求する上での課題を強調している。 - 具体的な課題の内容や限界については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている先行研究: estimation models と tactile feedback control laws を設計するアプローチ、tactile data から object-motion estimation・action-outcome prediction・action selection を学習するモデル。 - 同分野の関連手法として、general-purpose vision-language model (VLM) を用いたロボット操作や contact-rich manipulation の研究が挙げられる。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Kazutoshi Tanaka

分類: cs.RO

原文アブストラクト

Robots that autonomously determine their actions from language instructions and sensory observations could perform new contact-rich manipulation tasks without task-specific training or hand-designed rules. To perform these tasks, robots must infer how objects contact one another and move as a result, then select actions. For contact inference and action selection, prior approaches involve designing estimation models and tactile feedback control laws, or learning models for object-motion estimation, action-outcome prediction, and action selection from tactile data. Instead, we propose TacZero, which uses a pretrained general-purpose vision-language model (VLM) to interpret visual and tactile observations and select robot actions without additional tactile or manipulation training or task-specific rules for contact interpretation or action selection. TacZero provides the VLM with camera images, robot state, and three-axis tactile responses represented as numerical values or vectors overlaid on the images. From these observations and interaction history, the VLM generates commands specifying target end-effector positions and gripper opening or closing, which a low-level controller executes. In real-world cylindrical-peg insertion experiments, TacZero succeeded in 15 of 20 trials with numerical tactile input, compared with 10 of 20 without tactile input. This study provides a concrete starting point for further research on contact-rich manipulation using general-purpose VLMs and highlights challenges in pursuing this direction.

関連論文

PR本紙発行元 EmplifAI