日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLA/ヒューマノイドarXiv:2608.17453v1

EATR-Stereo: 身体性を考慮したステレオ証拠のルーティングによるヒューマノイド視覚言語行動制御

EATR-Stereo: Embodiment-Aware Routing of Paired Stereo Evidence for Humanoid Vision-Language-Action Control

シェア:XThreadsFacebookLINEはてブBluesky

頭部搭載ステレオカメラを持つヒューマノイドのVLA制御において、主視点のトークンを保持しつつ補助視点の情報を身体状態に応じて選択的に統合するフレームワークを提案し、実機で高い成功率を達成した。

詳しい要約

1. どんなもの?

EATR-Stereoは、頭部搭載ステレオカメラを用いた長期的なヒューマノイドVLA制御のための、embodiment-awareなトークンルーティングフレームワークである。プライマリビューのトークンを保持しつつ、同期された補助ビューのトークン列をクエリしてプライマリに整合したCross-View Auxiliary Tokens (CVATs)を構築する。さらに、身体セグメント化されたproprioceptiveエンコーダが、ロボットの構成履歴に基づいてトークンごとの補助情報の使用を調整し、行動生成中にステレオ証拠を選択的に組み込む。事前学習済みVLAの言語・視覚コンテキストを拡張しつつ、そのvision-languageモデルは凍結したままである。

2. 先行研究と比べてどこがすごい?

既存のインターフェースは、補完的なステレオ証拠を破棄するか、追加の観測を融合する際にネイティブなプライマリビューパスを保持せず、補助情報をロボットのembodimentに適応させない。EATR-Stereoは、プライマリビューのトークンを保持し、補助情報をembodimentに適応させる点で優れている。具体的には、proprioceptive状態履歴を用いてトークンごとの補助使用を調整することで、ステレオ証拠の選択的組み込みを実現し、事前学習済み表現との互換性を維持する。

3. 技術・手法の肝は?

手法の核心は、embodiment-awareなトークンルーティングである。まず、プライマリビューのトークンをそのまま保持する。次に、同期された補助ビューのトークン列をクエリして、プライマリに整合したCVATsを生成する。さらに、身体セグメント化されたproprioceptiveエンコーダが、ロボットの構成履歴を処理し、トークンごとの補助情報の使用を調整する。このルーティングされた補助ストリームは、事前学習済みVLAの言語・視覚コンテキストを拡張し、vision-languageモデルは凍結されたままである。

4. どうやって有効だと検証した?

33-DoFの物理ヒューマノイドと37次元のproprioceptive状態を用いて、100回以上の探索・接近・把持・配置・戻るタスクで9つの構成を評価した。EATR-Stereoは、フルタスク成功率60.0%、把持成功率100.0%、ステージ成功率80.0%を達成した。重度の非対称遮蔽下では、CVAT単独の30%に対して回復率80%を達成した。アブレーション研究により、プライマリトークンの保持と、クロスビュー補助特徴と構造化されたproprioceptiveルーティングの組み合わせの重要性が示された。

5. 議論はある?

要旨からは、議論の詳細は不明である。ただし、結果は選択的にルーティングされたペアのステレオ証拠が、長期的なヒューマノイドVLA制御の空間的接地を改善することを示している。アブレーション研究は、プライマリトークンの保持とproprioceptiveルーティングの重要性を強調しているが、限界や将来の方向性については言及されていない。

6. 次に読むべき論文は?

要旨で参照されている関連研究は明示されていないが、同分野の定番として、Vision-Language-Actionモデル(例:RT-2、OpenVLA)や、ステレオ視覚を用いたロボット制御、proprioceptive情報の統合に関する研究が考えられる。具体的には、OpenVLAやRT-2などのVLAモデル、およびステレオマッチングやクロスビュー特徴融合に関する論文が関連する。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Songwei Wu, Rui Zhao, Fan Yang, Zhongqiang Nie, Zhiduo Jiang, Wandong Sun, Yuwei Li, Yang Liu, Hong Liu

分類: cs.RO

原文アブストラクト

Long-horizon humanoid vision--language--action (VLA) control with head-mounted stereo cameras requires visual interfaces that can exploit complementary views while maintaining compatibility with pretrained representations. Existing interfaces often discard complementary stereo evidence or fuse additional observations without preserving the native primary-view pathway and adapting auxiliary information to robot embodiment. We present EATR-Stereo, an embodiment-aware token-routing framework that retains primary-view tokens and constructs primary-aligned Cross-View Auxiliary Tokens (CVATs) by querying the synchronized auxiliary-view token sequence. A body-segmented proprioceptive encoder further conditions token-wise auxiliary usage on robot configuration history, enabling selective incorporation of stereo evidence during action generation. The routed auxiliary stream augments the language and primary-visual context of a pretrained VLA while keeping its vision--language model frozen. On a 33-DoF physical humanoid with a 37-D proprioceptive state, we evaluate nine configurations in over-100-s search--approach--grasp--place--return tasks. EATR-Stereo achieves 60.0% full-task success, 100.0% grasp success, and 80.0% stage success. Under severe asymmetric occlusion, it improves recovery to 80% compared with 30% for CVAT alone. Ablation studies further show the importance of preserving primary tokens and combining cross-view auxiliary features with structured proprioceptive routing. These results demonstrate that selectively routed paired stereo evidence improves spatial grounding for reliable long-horizon humanoid VLA control.