日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.09734

ΔWAM: 行動接ベクトル場を世界行動モデルへ蒸留する手法

ΔWAM: Distilling Action Tangent Fields into World Action Models

シェア:XThreadsFacebookLINEはてブBluesky

行動に依存する未来変化を局所テイラー展開で抽出し、それを世界行動モデルに蒸留することで、環境変化に対するロボット方策の頑健性を高めた。

詳しい要約

1. どんなもの?

- ロボット政策を改善する World Action Models (WAM) の新しい学習手法 ΔWAM を提案。 - 未来予測の監督信号を、行動に関連する変化に集中させる Action Tangent Fields を導入。 - 外観やシーン持続性に支配されがちな未来予測を、行動依存のダイナミクスに焦点化。 - LIBERO-Plus, RoboTwin, RoboTwin2.0-Plus で環境摂動に対する頑健性を向上。 - 大規模 embodied pretraining なしで、いくつかの分布シフトで pretrained 政策より強い頑健性を達成。

2. 先行研究と比べてどこがすごい?

- 従来の WAM は optical flow, motion-centric representations, latent actions などで未来監督を効率化。 - しかし予測可能な未来の多くは外観やシーン持続性に支配され、行動依存の変化が少ない。 - 本研究は「行動関連変動の割合が高いほど世界監督が効率的」という共通視点を提示。 - Action Tangent Fields により、単に予測可能な外観ではなく行動と強く結合したダイナミクスへ監督を誘導。 - 大規模事前学習なしで、いくつかの分布シフト下で pretrained 政策を上回る頑健性を実現。

3. 技術・手法の肝は?

- 未来ダイナミクスを Residual-VAE 空間で表現。未来潜在は現在潜在とその残差から復元可能。 - 強い action-conditioned world model (ACWM) を用い、行動変動と残差世界変動の局所対応を調査。 - 行動が未来ダイナミクスに与える変化を局所テイラー展開で再定式化し、Action Tangent Fields を構築。 - この局所一次構造を WAM に蒸留し、デノイジング監督を行動と密結合したダイナミクスへ誘導。 - 多段 VideoDiT デノイジングを1ステップに蒸留し、効率的な推論を実現。

4. どうやって有効だと検証した?

- LIBERO-Plus, RoboTwin, RoboTwin2.0-Plus の3ベンチマークで評価。 - 照明、背景、カメラ、レイアウトなどの環境摂動に対する頑健性を検証。 - 大規模 embodied pretraining なしで、いくつかの分布シフト下で pretrained 政策より強い頑健性を確認。 - 多段 VideoDiT デノイジングの1ステップ蒸留による効率的推論も検証。

5. 議論はある?

- 有効な WAM 監督は情報豊富でありつつ、行動が未来を変える方向に予測能力を集中すべきと示唆。 - 行動関連変動の割合を高めることが世界監督の効率化に重要という洞察を提示。 - 限界や失敗事例、計算コスト、スケーラビリティに関する議論は要旨からは不明。

6. 次に読むべき論文は?

- optical flow を用いた WAM 研究 - motion-centric representations を用いた WAM 研究 - latent actions を用いた WAM 研究 - action-conditioned world model (ACWM) - Residual-VAE - VideoDiT - LIBERO-Plus, RoboTwin, RoboTwin2.0-Plus の関連研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ke Wu, Hanwen Huang, Bo Gu, Kaizhao Zhang, Xiangting Meng, Yupeng Zheng, Zijun Xu, Jieru Zhao, Wenchao Ding

分類: cs.RO

原文アブストラクト

World Action Models (WAM) improve robot policies by augmenting sparse action supervision with dense future prediction. However, much of the predictable future is dominated by appearance and scene persistence rather than action-dependent dynamics. We observe that several recent WAM designs, including optical flow, motion-centric representations, and latent actions, can be understood from a common perspective in which world supervision becomes more efficient as it contains a higher proportion of action-relevant variation. Based on this insight, we introduce Action Tangent Fields, which reformulate world supervision through a local Taylor expansion of how actions induce changes in future dynamics. We represent future dynamics in Residual-VAE space, where the future latent remains recoverable from the current latent and its residual, and use a strong action-conditioned world model (ACWM) to probe the local correspondence between action variations and residual-world variations. This local first-order structure is distilled into the WAM to guide its denoising supervision toward dynamics that are more tightly coupled to action, rather than merely predictable from appearance. Across LIBERO-Plus, RoboTwin, and RoboTwin2.0-Plus, our method consistently improves robustness to lighting, background, camera, layout, and other environmental perturbations. Despite using no large-scale embodied pretraining, it achieves stronger robustness under several distribution shifts than pretrained policies. We further distill multi-step VideoDiT denoising into a single step for efficient inference. Our results suggest that effective WAM supervision should remain information-rich while concentrating its predictive capacity on the directions along which actions change the future.

関連論文

PR本紙発行元 EmplifAI