日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.38163

世界行動モデルの表現設計を見直すReWAM

Rethinking Representations for World-Action Modeling

シェア:XThreadsFacebookLINEはてブBluesky

事前学習済みDINO特徴を基盤に、特徴校正と時間表現ボトルネックで世界状態を圧縮し、行動損失のみで表現を整形する世界行動モデルを提案。動画生成事前学習なしでRoboTwin 2.0で93.6%の成功率を達成。

詳しい要約

1. どんなもの?

- 世界行動モデル(World-Action Model)の表現空間設計を研究 - 制御と予測のインターフェースとしての表現を制御比較で検討 - 再構成忠実度や事前学習知覚特徴だけでは効果的な政策学習を保証しないことを発見 - この知見に基づき、事前学習済みDINO特徴を用いたReWAMを提案 - 表現中心の世界行動モデルで、Feature CalibrationとTemporal Representation Bottleneckで特徴を整理 - Action-Grounded Representation Shapingで行動損失勾配のみをボトルネックに流す - 生成ビデオ事前学習なしでRoboTwin 2.0で93.6%成功、RoboDojoで平均スコア12.29、成功率8.28%を達成

2. 先行研究と比べてどこがすごい?

- 従来の世界行動モデルは再構成忠実度や事前学習知覚特徴に依存 - 本研究はそれらだけでは不十分であることを制御比較で示す - 生成ビデオ事前学習なしで高い性能を達成 - 表現空間の設計を体系的に検討した点が新しい - 行動損失勾配のみをボトルネックに流すことで政策が表現を形成 - 世界モデルは表現の進化を学習する分離アーキテクチャ

3. 技術・手法の肝は?

- 事前学習済みDINO特徴を基盤に使用 - Feature Calibrationで特徴を調整 - Temporal Representation Bottleneckでコンパクトな世界状態を形成 - Action-Grounded Representation Shapingで行動損失勾配のみをボトルネックに伝播 - 政策が表現内容を形成し、世界モデルがその進化を学習 - 生成ビデオ事前学習を必要としない

4. どうやって有効だと検証した?

- RoboTwin 2.0で93.6%の成功率を達成 - RoboDojoで平均スコア12.29、成功率8.28%を達成 - 約600時間の身体性事前学習データを使用 - 制御比較により表現設計の影響を検証 - 再構成忠実度と事前学習知覚特徴の限界を実験的に示す

5. 議論はある?

- 表現空間の設計が政策学習と予測に与える影響を議論 - 再構成忠実度や事前学習特徴だけでは不十分である理由を考察 - 行動損失勾配のみを流すことの有効性と限界 - 生成ビデオ事前学習なしでの性能達成の意義 - 身体性事前学習データの規模と性能の関係 - 要旨からは詳細な議論の内容は不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない - 関連手法としてWorld-Action Model、DINO、Feature Calibration、Temporal Representation Bottleneck、Action-Grounded Representation Shaping - 同分野の定番としてRoboTwin 2.0、RoboDojo - 次に読むべき論文は要旨からは特定できない

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Haoyi Jiang, Liu Liu, Xinjiang Wang, Zhihao Sun, Zequn Chen, Sen Wang, Xinjie Wang, Xia Chen, Jingfeng Yao, Weiheng Zhao, Shanglin Yuan, Zhizhong Su, Wei Sui, Wenyu Liu, Xinggang Wang

分類: cs.CV, cs.RO

原文アブストラクト

World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure effective policy learning. These findings motivate ReWAM, a representation-centric world-action model built on pre-trained DINO features. Feature Calibration and a Temporal Representation Bottleneck organize these features into compact world states suited to dynamics modeling. Action-Grounded Representation Shaping routes only action-loss gradients to the bottleneck, thereby letting the policy shape what the representation encodes while the world model learns how it evolves. Without generative video pre-training, ReWAM achieves 93.6% success on RoboTwin 2.0. On RoboDojo, it achieves an average score of 12.29 and a success rate of 8.28% using approximately 600 hours of embodied pre-training data.

関連論文

PR本紙発行元 EmplifAI