日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.18497

TAO-Force: 接触の多いマニピュレーションのための力認識知覚と高速・低速制御の統合

TAO-Force: Unifying Force-Aware Perception and Fast-Slow Control for Contact-Rich Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

力フィードバックを視覚言語モデルに注入し、接触時のみ高速なアドミタンス制御に切り替える枠組みを提案し、接触の多い操作タスクでの有効性を示した。

詳しい要約

1. どんなもの?

- 視覚中心のVLAモデルは接触を伴う操作に不十分。 - 力覚フィードバックを統合したTAO-Forceを提案。 - 力条件付きVLAフレームワークで、力認識知覚と接触調整実行を統合。 - 接触リッチな操作タスク向け。

2. 先行研究と比べてどこがすごい?

- 従来のVLAは視覚中心で位置制御のみ。 - 接触開始や相互作用の大きさを視覚だけでは捉えにくい。 - 位置制御では急変する接触ダイナミクスに柔軟に対応できない。 - TAO-Forceは力覚を統合し、接触フェーズに応じた制御を実現。

3. 技術・手法の肝は?

- Force-conditioned Feature-wise Linear Modulation (F-FiLM)を導入。 - 凍結した事前学習済み視覚言語バックボーンの表現に力覚フィードバックを注入。 - 意味的事前知識を保持。 - 接触ゲート付きfast-slowアーキテクチャを採用。 - 非接触時はslow位置制御ブランチが公称軌道を追従。 - 接触時はfastアドミタンス制御ブランチが物理相互作用を調整。

4. どうやって有効だと検証した?

- 力知覚タスクの詳細分析を実施。 - 実世界で4つの接触リッチ操作タスクを評価。 - 有効性とロバスト性を検証。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明記されていない。 - 関連手法としてVision-Language-Action (VLA)モデル、アドミタンス制御、Feature-wise Linear Modulation (FiLM)が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Bohan Gan, Xuanzhang Wen, Yongsheng Zhao, Baoping Cheng, Wenhe Jia, Ye Wang, Gongxin Yao, Han Gao, Jingyao Tang, Lei Zhao, Ji Ge

分類: cs.RO

原文アブストラクト

Vision-Language-Action (VLA) models have demonstrated strong performance across diverse robotic manipulation tasks, yet their predominantly vision-centric perception and position-controlled execution remain insufficient for contact-rich manipulation. Visual observations alone often provide limited evidence of contact onset and interaction magnitude, while position-control policies cannot respond compliantly to rapidly changing contact dynamics. To bridge both the perception and control gaps, we propose TAO-Force, a force-conditioned VLA framework that combines force-aware policy learning with contact-regulated execution. For force-aware perception, TAO-Force introduces Force-conditioned Feature-wise Linear Modulation (F-FiLM) to inject encoded force feedback into the representations of a frozen pretrained visual-language backbone while preserving its semantic priors. For responsive control, it employs a contact-gated fast-slow architecture, with a slow position-control branch tracking nominal trajectories during non-contact phases and a fast admittance-control branch regulating physical interaction during contact phases. Detailed analyses on a force-perception task and real-world evaluations across four contact-rich manipulation tasks validate the effectiveness and robustness of TAO-Force.

関連論文

PR本紙発行元 EmplifAI