日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.12285

PLaW-VLA: 予測的潜在世界モデリングによる視覚-言語-行動ポリシー

PLaW-VLA: Predictive Latent World Modeling for Vision-Language-Action Policies

シェア:XThreadsFacebookLINEはてブBluesky

事前学習済みの予測指向表現空間でタスクに関連する将来状態をモデル化し、長期制御と分布シフトへの汎化を改善するVLAポリシーを提案。

詳しい要約

1. どんなもの?

- PLaW-VLAは、Vision-Language-Action (VLA) ポリシーのための予測的潜在世界モデリング手法である。 - タスクに関連する未来状態を、事前学習済みの予測指向表現空間でモデル化する。 - これにより、制御に無関係な視覚的詳細の予測を減らし、長期的制御のための予測コンテキストを提供する。 - Mixture-of-Transformersアーキテクチャに基づき、観察履歴、現在のタスク意味論、予測未来状態を構造的因果注意で条件付ける。

2. 先行研究と比べてどこがすごい?

- 反応的ポリシーと比較して、RoboTwin Hard Horizon IIIで+11.8パーセントポイント (pp) の性能向上。 - 再構成指向の潜在予測と比較して、ゼロショットLIBERO-Plusで+1.77 ppの向上。 - 低レベルの視覚再構成を避けることで、未来予測の負担を軽減し、軽量な潜在世界モデルを実現。 - 生成的世界行動モデリングと比較して、同程度のポリシー性能で推論レイテンシが約1/19。

3. 技術・手法の肝は?

- 事前学習済みの予測指向表現空間でタスク関連未来状態をモデル化する。 - Mixture-of-Transformersアーキテクチャを採用。 - 観察履歴、現在のタスク意味論、予測未来状態を構造的因果注意で条件付けて行動生成を行う。 - 並列未来予測を可能にする軽量な潜在世界モデルを構築。

4. どうやって有効だと検証した?

- RoboTwin Hard Horizon IIIでの実験により、反応的ポリシーに対する+11.8 ppの性能向上を確認。 - ゼロショットLIBERO-Plusでの実験により、再構成指向の潜在予測に対する+1.77 ppの向上を確認。 - これらの結果から、長期的制御の改善と分布シフト下での汎化性能の向上を検証。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: 反応的ポリシー、再構成指向の潜在予測、生成的世界行動モデリング。 - 関連手法: Vision-Language-Action (VLA) ポリシー、Mixture-of-Transformers、潜在世界モデル。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yu Liu, Hetian Guo, Tianlv Huang, Ziyi Cai, Wudi Chen, Hantang Wang, Qiutong Liu, Yingzhi Peng, Wei Han, Peijun Tang, Jianan Wang, Zipei Fan, Zhiyuan Zha, Xuan Song

分類: cs.RO

原文アブストラクト

Learning to predict how the world evolves can provide vision-language-action (VLA) policies with predictive context for long-horizon control, but its effectiveness depends on what future representation is modeled and how it conditions action generation. We introduce PLaW-VLA, which models task-relevant future states in a pretrained prediction-oriented representation space, reducing the need to predict control-irrelevant visual details. Built on a Mixture-of-Transformers architecture, PLaW-VLA conditions action generation on observation history, current task semantics, and predicted future states through structured causal attention. Experiments show a +11.8 percentage-point (pp) gain over reactive policies on RoboTwin Hard Horizon III and a +1.77 pp gain over reconstruction-oriented latent prediction on zero-shot LIBERO-Plus, supporting improved long-horizon control and generalization under distribution shift, respectively. By avoiding low-level visual reconstruction, PLaW-VLA lowers the burden of future prediction, enabling a lightweight latent world model with parallel future prediction and about 1/19 the inference latency of generative world-action modeling at comparable policy performance.

関連論文

PR本紙発行元 EmplifAI