日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.25785

VisForce: 目標条件付き巧みな操作のための現在力と目標力の視覚的接地

VisForce: Visual Grounding of Current and Desired Forces for Goal-Conditioned Dexterous Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

指先位置に現在力と目標力を視覚的に重畳し、目標条件付きクロスアテンションで力認識動作を生成するVLA手法を提案し、実機で把持・操作タスクを評価した。

詳しい要約

1. どんなもの?

- Vision-Language-Action (VLA) モデルを用いた巧みな手の操作において、接触力を視覚的に接地する VisForce を提案。 - 現在の力と目標の力を指先位置に対応付けて視覚的手がかりとして描画。 - 手首画像とタスク固有の目標画像に力をレンダリングし、goal-conditioned cross-attention で統合。 - 力に敏感な行動を生成する。

2. 先行研究と比べてどこがすごい?

- 従来の VLA では接触力を別状態や力特化表現として扱い、力と視覚位置の空間対応が明示しにくかった。 - VisForce は力を指先位置に視覚的に接地することで、この対応を明示的に表現。 - 力対応の条件付けを VLA ベースの巧みな手の操作に有効に用いる点が新しい。

3. 技術・手法の肝は?

- 現在の力と目標の力を、現在の手首画像とタスク固有の目標画像上に視覚的力キューとしてレンダリング。 - 2つの表現を goal-conditioned cross-attention で統合。 - 統合表現から力対応行動を生成。 - 指先位置に整列した視覚的力表現が肝。

4. どうやって有効だと検証した?

- 実機 UR10 ロボットと RH56F1 巧みな手を使用。 - 力条件付け把持と3つの多段階操作タスクで評価。 - 力条件付け把持では、目標力の増加に伴い一貫した把持力応答。 - 卵と歯磨き粉チューブの把持・持ち上げ成功率はそれぞれ70%、80%。 - カップ挿入/ボトル注ぎ、トング支援パン移動、滑り調整ペグインホールの最終成功率はそれぞれ70%、55%、40%。

5. 議論はある?

- 指先に整列した視覚的力表現が、VLA ベースの巧みな手の操作における力対応条件付けに有効であることを示す。 - 限界や課題、今後の議論については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として Vision-Language-Action (VLA) モデル、goal-conditioned cross-attention が挙げられる。 - 同分野の定番として dexterous hand manipulation、force-aware manipulation に関する研究が考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jung-Woo Lee, Soo-Chul Lim

分類: cs.RO

原文アブストラクト

Vision-Language-Action (VLA) models have emerged as general-purpose robotic manipulation policies. However, in dexterous hand manipulation, contact forces are typically provided as separate states or force-specific representations, making it difficult to explicitly represent the spatial correspondence between force and their corresponding visual locations. In this work, we propose VisForce, which visually grounds the current and desired forces at their corresponding fingertip locations. VisForce renders current and desired visual force cues on the current wrist image and a task-specific goal image, and combines the two representations through goal-conditioned cross-attention to generate force-aware actions. We evaluate VisForce using a real UR10 robot equipped with an RH56F1 dexterous hand through force-conditioned grasping and three multi-stage manipulation tasks. In force-conditioned grasping experiments, VisForce exhibited a consistent grip-force response as the desired force increased, and achieved grasp-and-lift success rates of 70% and 80% for an egg and a toothpaste tube, respectively. It further achieved final success rates of 70%, 55%, and 40% on cup insertion/bottle pouring, tong-assisted bread transfer, and slip-modulated peg-in-hole, respectively. These results show that fingertip-aligned visual force representations can be effectively used for force-aware conditioning in VLA-based dexterous hand manipulation.

関連論文

PR本紙発行元 EmplifAI