日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.02054

UniWAM:統合された世界行動モデル

UniWAM: Unified World-Action Model

シェア:XThreadsFacebookLINEはてブBluesky

物理推論・映像生成・行動予測を統合したモデルで、人間の一人称視点データとロボットデータを活用し、複数の評価でSOTAを達成した。

詳しい要約

1. どんなもの?

- 統合アーキテクチャUniWAMを提案 - 物理推論器、世界生成器、行動予測器を統合 - 物理世界の意味理解、視覚生成、行動予測を同時学習 - 人間の一人称視点データとロボットデータの厳格なクリーニング・アノテーションパイプラインを開発 - 低レベル行動を自然言語で表現し、VQAデータ、人間一人称データ、ロボットデモンストレーションから補完的監督を割り当てる事前学習レシピを導入 - ポストトレーニングでは、未来視覚ノイズ拡張と履歴条件付きフローマッチングを採用 - 複数の評価でSOTA性能を達成し、人間-ロボット共同トレーニングの対数線形スケーリング則を発見

2. 先行研究と比べてどこがすごい?

- Vision-language-actionモデルは事前学習済み視覚言語モデルの理解・推論能力を活用するが、行動のみの監督では世界ダイナミクスへの接地が限定的 - World-actionモデルはビデオ生成モデルから時空間事前知識を継承するが、分布シフト下での意味理解・推論に制限 - UniWAMはこれらを統合し、意味理解、視覚生成、行動予測を共同学習することで、両者の利点を併せ持つ - 従来手法と比較して、分布内性能、ロバスト性、一般化、指示追従、長期タスク実行でSOTAを達成

3. 技術・手法の肝は?

- 物理推論器、世界生成器、行動予測器を統合した統一アーキテクチャ - 低レベル行動を自然言語で表現し、視覚言語コンポーネントを embodied タスクに適応させつつ言語能力を保持 - 事前学習レシピ:VQAデータ、人間一人称データ、ロボットデモンストレーションから適切なモデルコンポーネントに補完的監督を割り当て - ポストトレーニング:未来視覚ノイズ拡張により正確な未来予測への依存を低減、履歴条件付きフローマッチングでエンコードされた行動履歴を行動生成の初期化に使用 - これらの設計により、性能を維持しつつデノイジングステップを大幅に削減

4. どうやって有効だと検証した?

- 複数の評価でSOTA性能を達成 - 分布内性能、ロバスト性、一般化、指示追従、長期タスク実行 - 人間-ロボット共同トレーニングの対数線形スケーリング則を発見し、大規模事前学習の有効性を実証 - 具体的なデータセット名やベンチマーク名は要旨からは不明

5. 議論はある?

- 人間-ロボット共同トレーニングの対数線形スケーリング則を発見 - 大規模事前学習が人間とロボットの混合データで有効であることを示唆 - その他の議論や限界については要旨からは不明

6. 次に読むべき論文は?

- Vision-language-actionモデル(例:RT-2, OpenVLAなど) - World-actionモデル(例:UniPi, World Modelsなど) - ビデオ生成モデル(例:Video Diffusion Modelsなど) - フローマッチング(Flow Matching) - 具体的な論文名は要旨からは不明

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jiayi Chen, Wenxuan Song, Jingbo Wang, Shuai Zhou, Xicheng Gong, Zehua Fan, Ziyang Zhou, Junwu E, Haodong Yan, Fuhao Li, Qize Yu, Xu Huang, Pengwei Wang, Wen Chen, Shunbo Zhou, Haoang Li

分類: cs.RO

原文アブストラクト

Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal priors from video generation models, yet remain limited in semantic understanding and reasoning under distribution shifts. We introduce UniWAM, a unified architecture that integrates a physical reasoner, a world generator, and an action predictor to jointly learn semantic understanding of the physical world, visual generation, and action prediction. To ensure the quality of the training data, we developed a rigorous data cleaning and annotation pipeline for both human egocentric data and robot data. To adapt the vision-language component to embodied tasks while preserving its inherited language capabilities, we represent low-level actions in natural language and introduce a pre-training recipe that assigns complementary supervision from visual question answering (VQA) data, human egocentric data, and robot demonstrations to the appropriate model components. During post-training, future visual noise augmentation reduces reliance on precise future predictions, while history-conditioned flow matching uses encoded action history to initialize action generation. Together, these designs significantly reduce denoising steps while maintaining performance. UniWAM achieves state-of-the-art (SOTA) performance across multiple evaluations, including in-distribution performance, robustness, generalization, instruction following, and long-horizon task execution. Furthermore, we uncover a log-linear scaling law of unified human-robot co-training, demonstrating the effectiveness of large-scale pre-training on a mixture of human and robot data.

関連論文

PR本紙発行元 EmplifAI