日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.12194

MiniWAM: 効率的なワールド・アクション・モデリングのためのコンパクトな未来ターゲットの学習

MiniWAM: Learning Compact Future Targets for Efficient World-Action Modeling

シェア:XThreadsFacebookLINEはてブBluesky

未来状態を高次元のまま予測する従来のWorld Action Modelに対し、逆時空間モデリングで制御に関連するコンパクトな未来表現を学習し、予測コストを大幅に削減しつつ性能を向上させた手法。

詳しい要約

1. どんなもの?

- World Action Models (WAMs) の一種で、ロボット方策の co-training 目的として未来状態と行動を同時予測する枠組み。 - 従来の WAMs は pretrained visual backbones の native 表現空間で未来状態を予測し、高次元ターゲットで学習コストが大きい。 - MiniWAM は privileged な current-future transitions から学習した compact な未来表現を予測対象とする。 - 0.25B パラメータで LIBERO, LIBERO-Plus, RoboTwin 2.0 の simulation benchmarks において大規模 WAMs と競合。

2. 先行研究と比べてどこがすごい?

- 従来の WAMs は native future feature を予測し、高次元で学習コスト大。 - MiniWAM は native future feature tokens を 65× 削減。 - DINOv3 と WAN2.1 VAE 特徴の両方で native future-feature prediction を一貫して上回る。 - world-action training で最大 8× の高速化を達成。 - 0.25B パラメータで、LIBERO, LIBERO-Plus, RoboTwin 2.0 において大幅に大きい WAMs と競合。

3. 技術・手法の肝は?

- Predictive Representations via Inverse Spatiotemporal Modeling (PRISM) を提案。 - PRISM は inverse-dynamics supervision と feature reconstruction を組み合わせ、制御に関連する遷移情報を強調しつつ有用な未来状態情報を保持。 - 学習した PRISM encoder を frozen にし、MiniWAM は現在観測から得られるターゲットとロボット行動を同時予測するよう訓練。 - compact な未来表現をターゲットとすることで効率化。

4. どうやって有効だと検証した?

- LIBERO, LIBERO-Plus, RoboTwin 2.0 の simulation benchmarks で評価。 - DINOv3 と WAN2.1 VAE 特徴の両方で native future-feature prediction と比較し、一貫した性能向上を確認。 - world-action training で最大 8× の speedup を検証。 - 0.25B パラメータで大規模 WAMs と競合することを確認。 - 表現分析により、PRISM が feature reconstruction 単独を超える behavioral structure に寄与することを示す。

5. 議論はある?

- 有効な world-action modeling には native visual futures の予測は不要であることを示唆。 - compact predictive representations が強力かつ大幅に効率的な policy learning のターゲットを提供。 - 具体的な限界や失敗ケース、計算資源の詳細、実機実験の有無については要旨からは不明。

6. 次に読むべき論文は?

- DINOv3 - WAN2.1 VAE - LIBERO - LIBERO-Plus - RoboTwin 2.0 - World Action Models (WAMs) 関連の先行研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jie Chen, Ruofei Bai, Yuxin Cai, Yifeng Zhang, Chengyang He, Jun Li, Wei-Yun Yau, Guillaume Sartoretti

分類: cs.RO

原文アブストラクト

World modeling has emerged as an effective co-training objective for robot policies, giving rise to World Action Models (WAMs) that jointly predict actions and future states. However, most WAMs predict future states in the native representation space of pretrained visual backbones, resulting in high-dimensional targets with substantial training cost. We introduce MiniWAM, which instead predicts compact future representations learned from privileged current-future transitions. To construct these targets, we propose Predictive Representations via Inverse Spatiotemporal Modeling (PRISM), which combines inverse-dynamics supervision with feature reconstruction to emphasize control-relevant transition information while preserving useful future-state information. With the learned PRISM encoder frozen, MiniWAM is trained to jointly predict the resulting targets and robot actions from current observations. With 65$\times$ fewer native future feature tokens, MiniWAM consistently outperforms native future-feature prediction with both DINOv3 and WAN2.1 VAE features, while achieving up to an 8$\times$ speedup in world-action training. At 0.25B parameters, MiniWAM is already competitive with substantially larger WAMs on LIBERO, LIBERO-Plus, and RoboTwin 2.0 simulation benchmarks. Representation analyses further show that PRISM contributes behavioral structure beyond feature reconstruction alone. These results demonstrate that effective world-action modeling does not require predicting native visual futures, and that compact predictive representations provide a strong and substantially more efficient target for policy learning. The project page is available at: https://j1dan.github.io/MiniWAM.

関連論文

PR本紙発行元 EmplifAI