日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2608.24882

潜在アクションを意図として用いた効率的な未来想像による世界行動モデル

Latent Action as Intention Enables Efficient Future Imagination for World Action Models

シェア:XThreadsFacebookLINEはてブBluesky

ロボット制御のための世界行動モデルにおいて、テスト時の未来観測生成の遅延を削減するため、潜在アクションを意図の表現として用いるLAWAを提案。将来の観測を生成せずに効率的な未来想像を実現し、少ないデモデータや分布外シナリオでの性能を向上させた。

詳しい要約

1. どんなもの?

LAWAは、World Action Models (WAMs)の推論時レイテンシ問題を解決する新しいアーキテクチャ。将来の観測を生成せずに、コンパクトな潜在アクションを将来の意図の表現として用いることで、効率的な将来想像を実現する。

2. 先行研究と比べてどこがすごい?

従来のWAMは将来観測の生成に高いレイテンシを要する。Fast-WAMはこれを除去するが、ロボットデモが少ない場合やOODシナリオで汎化性能が低下する。LAWAは将来情報を潜在アクションに圧縮することで、性能を保ちつつレイテンシを削減する。

3. 技術・手法の肝は?

離散トークナイザをアクションフリー事前学習で強化し、操作中心のコードブックターゲットを生成。LAWAはこのターゲットに基づく連続潜在状態と実行可能なアクションチャンクを共同でデノイズし、推論時には将来ビデオ分岐を省略する。

4. どうやって有効だと検証した?

RoboCasaで評価し、数ショット設定で平均成功率65.6%、フルデータ設定で80.8%を達成。Fast-WAMベースラインよりそれぞれ9.6ポイント、4.5ポイント向上。Joint-WAMと同等性能で、推論レイテンシを42.9%削減。LIBERO-Plusでゼロショット堅牢性、実世界タスクでも優れた性能を示した。

5. 議論はある?

要旨からは、潜在アクションの解釈可能性や、より複雑なタスクでのスケーラビリティ、他のモダリティへの適用可能性などは不明。また、Fast-WAMとの性能差の原因分析や、潜在アクションの品質が性能に与える影響についての詳細な議論は要旨に含まれていない。

6. 次に読むべき論文は?

要旨で参照されているFast-WAMとJoint-WAMの論文。また、関連するWorld Action Modelsの基礎研究や、アクションフリー事前学習、離散トークナイザの手法に関する論文が考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xiang Li, Yupeng Zheng, Songen Gu, Huailiang Ma, Feng Yu, Xian Nie, Shanshuai Yuan, Yujie Zang, Weize Li, Shuai Tian, Moyang Liu, Ya-Qin Zhang, Wenchao Ding

分類: cs.RO

原文アブストラクト

World action models (WAMs) improve robot control by modeling how observations evolve, but generating future observations at test time incurs substantial latency. Fast-WAM removes this process for efficiency; however, our matched implementations show lower generalization for Fast-WAM than for future-aware alternatives, especially with scarce robot demonstrations and in out-of-distribution scenarios. To bridge this gap, we introduce **LAWA**, a WAM architecture that uses compact latent actions as an operational representation of future intentions, enabling efficient test-time future imagination without generating future observations. Specifically, a discrete tokenizer enhanced by action-free pre-training produces manipulation-centric codebook targets. LAWA jointly denoises a continuous latent state anchored to these targets with executable action chunks while omitting the future-video branch at inference. On RoboCasa, LAWA achieves state-of-the-art average success rates of 65.6% and 80.8% in the few-shot and full data settings, improving over the matched Fast-WAM baseline by 9.6 and 4.5 points, respectively. It also preserves the performance level of the matched Joint-WAM variant while requiring 42.9% lower inference latency. LAWA also demonstrates competitive zero-shot robustness on LIBERO-Plus and superior performance on real-world tasks. These results show that future imagination need not be discarded: retaining it with compact latent actions yields an effective trade-off among performance, generalization, and latency. Code and models will be released.

関連論文