日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLA/移動操作arXiv:2608.02257v1

パノラマ認識型VLAによる全身遠隔操作を用いた移動操作の学習

Learning Panorama-Aware VLA for Mobile Manipulation with Whole-Body Teleoperation

シェア:XThreadsFacebookLINEはてブBluesky

全身遠隔操作システムとパノラマ認識型VLAポリシーを開発し、移動操作タスクの成功率を大幅に向上させた。

詳しい要約

1. どんなもの?

本論文は、移動操作ロボットのためのパノラマ認識型Vision-Language-Action (VLA)ポリシーであるPanoVLAを提案する。全身テレオペレーションシステムを用いて収集した5.5時間の実世界マルチモーダルデータセットを基に、パノラマ観測を統合するアーキテクチャを導入し、移動ベースと双腕の協調制御を実現する。

2. 先行研究と比べてどこがすごい?

既存のVLAモデルは主にローカルカメラ観測に依存し、視野が狭くグローバルな空間理解が困難である。本手法は、パノラマエンコーディングと融合モジュールを導入し、パノラマ観測と言語命令、ロボット状態を統合することで、この問題を解決する。また、全身テレオペレーションシステムにより、高品質な全身デモンストレーションの効率的な収集を可能にしている点が新しい。

3. 技術・手法の肝は?

PanoVLAはMixture-of-Transformersアーキテクチャに基づく。専用のパノラマエンコーディングと融合モジュールにより、パノラマ観測からグローバルな空間コンテキストを抽出し、言語命令やロボット状態と効果的に統合してアクション生成を行う。テレオペレーションシステムは、単一のVRインターフェースで車輪付き双腕ロボットの協調制御を可能にする。

4. どうやって有効だと検証した?

実世界の4つの移動操作タスクで評価し、平均ステージ完了率91.3%、エンドツーエンド成功率73.4%を達成し、ローカルビューのベースラインを大幅に上回った。これにより、パノラマ空間コンテキストの統合が空間理解と閉ループ操作性能を向上させることを実証した。

5. 議論はある?

要旨からは、パノラマ観測の導入が有効であることが示されたが、データセットの規模やタスクの多様性、他のセンサモダリティとの比較、計算コストなどに関する議論は明示されていない。また、テレオペレーションシステムのユーザビリティや一般化性能についての詳細は不明。

6. 次に読むべき論文は?

要旨で参照されている関連研究として、既存のVLAモデル(例:RT-2、Octo)や、移動操作のためのデータ収集手法(例:Open X-Embodiment)が挙げられる。また、パノラマ認識のための視覚エンコーダ(例:CLIP)や、Mixture-of-Transformersアーキテクチャの関連研究も参考になる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Donglin Yang, Haoran Chen, Xingyu Chen, Lixing Liu, Manyi Li, Changhe Tu, Ke Xu, Xiaojian Ma, Si Liu

分類: cs.RO

原文アブストラクト

Mobile manipulation is a key capability for embodied intelligence, enabling robots to accomplish complex multi-stage tasks in open-world environments. However, mobile manipulation poses two key challenges for vision-language-action (VLA) policies: At the data level, the efficient collection of high-quality whole-body demonstrations demands the coordinated control of both the mobile base and the robotic arms; at the model level, existing VLA models predominantly rely on local camera observations, whose limited field of view hinders global spatial understanding. To address these challenges, we develop a whole-body teleoperation system and a panoramic-aware VLA policy. The system enables coordinated control of a wheeled bimanual robot through a single VR interface and supports the acquisition of a real-world mobile manipulation dataset comprising 5.5 hours of multimodal demonstrations. Building upon this dataset, we propose PanoVLA, a panorama-aware vision-language-action policy for mobile bimanual manipulation. Built upon a Mixture-of-Transformers architecture, PanoVLA introduces global spatial context through dedicated panorama encoding and fusion modules, enabling effective integration of panoramic observations with language instructions and robot states for action generation. Evaluation on four real-world mobile manipulation tasks demonstrates that PanoVLA achieves an average stage completion rate of 91.3\% and an end-to-end success rate of 73.4\%, substantially outperforming local-view baselines. These results demonstrate that incorporating panoramic spatial context improves spatial understanding and closed-loop manipulation performance in mobile robots.