日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.33197

TAO-DA: 自律操作に向けた双腕協調マニピュレーションのためのデュアルアーム視覚言語行動モデル

TAO-DA: Towards Autonomous Operation--A Dual-Arm Vision-Language-Action Model for Coordinated Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

共有VLMバックボーンに腕ごとに分離したエキスパート塔を組み合わせ、言語指示と視覚意味による二段階の意図ルーティングで双腕の干渉を解消するVLAモデルを提案。タスク進捗予測モジュールで協調スケジューリングも可能にした。

詳しい要約

1. どんなもの?

- 提案: TAO-DAは、Dual-Arm Vision-Language-Action (VLA) モデルで、協調操作を自律的に行うためのフレームワーク。 - 目的: 既存VLAモデルが二本の腕の状態と意図を分離できず、意図しない腕間干渉が生じる問題を解決。 - 構成: 共有Vision-Language Model (VLM) バックボーン上に、腕ごとに分離されたexpert towersを持つ対称的なDual-Arm Expert (DAE) アーキテクチャ。 - 特徴: 二段階のdual-arm intent routing scheme(言語指示と視覚意味によるルーティング)と、タスク進捗予測モジュールを備える。

2. 先行研究と比べてどこがすごい?

- 先行VLAモデルは、二本の腕の状態と意図を明示的に分離するメカニズムを欠き、腕間干渉で成功率が低下。 - 提案手法は、腕ごとのexpert towersと二段階ルーティングにより、意図の分離と干渉の解消を実現。 - さらに、タスク進捗予測モジュールを導入し、協調マルチロボットタスクのスケジューリングを支援。 - 単腕から双腕(およびその逆)への創発的なスキル汎化と、腕間の運動ドメインスキル転移の予備的証拠を示す。

3. 技術・手法の肝は?

- 共有VLMバックボーン上に、腕ごとに分離されたexpert towersを持つ対称的Dual-Arm Expert (DAE) アーキテクチャ。 - 二段階のdual-arm intent routing scheme: 第一段階では明示的な言語指示、第二段階では暗黙的な視覚意味に基づいてexpertをルーティング。 - 軽量なタスク進捗予測モジュール: pre-chunk temporal featuresと固有受容感覚・視覚観測の意味表現間のcross-attentionを活用し、フレームごとのタスク完了進捗を推定。 - このモジュールがタスク進捗の同期を可能にし、協調マルチロボットタスクのスケジューリングを支援。

4. どうやって有効だと検証した?

- 実験結果により、dual-arm intent routingの有効性と腕間干渉の分離を実証。 - 単腕から双腕(およびその逆)への創発的なスキル汎化の予備的証拠を提示。 - 腕間の運動ドメインスキル転移の予備的証拠も示す。 - 具体的なデータセットや評価指標は要旨からは不明。

5. 議論はある?

- 要旨からは、提案手法の限界や議論の詳細は不明。 - ただし、実験結果はdual-arm intent routingと腕間干渉の分離における有効性を示し、スキル汎化と転移の予備的証拠を提供。 - 協調マルチロボットタスクのスケジューリング支援についても言及。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、Vision-Language-Action (VLA) モデル、Dual-Arm Expert (DAE) アーキテクチャ、cross-attention、intent routing schemeが挙げられる。 - 同分野の定番として、RT-1、RT-2、Octo、OpenVLAなどのVLAモデルが考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yongsheng Zhao, Han Gao, Baoping Cheng, Jingyao Tang, Dian Zhou, Deng Liang, Ji Ge, Xuanzhang Wen, Lei Zhao, Ye Wang

分類: cs.RO, cs.AI

原文アブストラクト

Vision-Language-Action (VLA) models provide a unified framework for grounding high-level semantic information into low-level robot actions, enabling scalable robotic manipulation across diverse tasks. However, existing VLA models lack explicit mechanisms to disentangle the states and intents of the two arms, leading to unintended cross-arm interference that degrades task execution success. To address this issue, we propose a symmetric Dual-Arm Expert (DAE) architecture built upon a shared Vision-Language Model (VLM) backbone with decoupled, arm-specific expert towers. Expert selection is carried out through a two-stage dual-arm intent routing scheme, in which experts are routed either by explicit language instructions in the first stage or by implicit visual semantics in the second stage. Moreover, we introduce a lightweight task progress prediction module that leverages cross-attention between the pre-chunk temporal features and semantic representations of proprioceptive and visual observations to accurately estimate frame-wise task completion progress. This module facilitates task progress synchronization to support coordinated scheduling for collaborative multi-robot tasks. Experimental results demonstrate the effectiveness of our model in dual-arm intent routing and the disentanglement of cross-arm interference, and further provide preliminary evidence of emergent skill generalization from single- to dual-arm tasks (as well as the reverse), together with cross-arm motion-domain skill transfer.

関連論文

PR本紙発行元 EmplifAI