日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.00859

CtrlWAM: 意図と予見を整合させた制御可能なワールドアクションモデル

CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight

シェア:XThreadsFacebookLINEはてブBluesky

摂動させた行動をシミュレータで実行し、その視覚的結果と対応付けて学習することで、行動予測と映像生成の整合性を高めた制御可能なワールドアクションモデルを提案。

詳しい要約

1. どんなもの?

- 本論文は World Action Models (WAMs) の新手法 CtrlWAM を提案する。 - WAMs は行動 (intent) と視覚的未来 (foresight) を同時に予測するモデル。 - 標準学習は記録行動と動画に同時にノイズを加えるが、摂動行動が反事実的未来を示唆する一方、ノイズ付き動画は GT 記録に縛られる不整合が生じる。 - CtrlWAM は摂動行動をシミュレータで実行し、そのノイズ付き視覚結果とペアにして WAM を学習する。 - 動画と行動の異なるノイズ除去要件に対応する warped video-action noise schedules を導入。 - 行動インターフェースを ego-only 制御から可変数のエージェントストリームに拡張。 - 運転実験とロボティクス実験で有効性を検証。プロジェクトページ: https://ctrl-wam.github.io/

2. 先行研究と比べてどこがすごい?

- 標準的な WAM 学習は記録行動と動画に同時にノイズを加えるため、摂動行動とノイズ付き動画の間に不整合が生じる。 - 低ノイズ領域では、ノイズ付き未来フレームからシーン形状や動的挙動が明確に見えるにもかかわらず、摂動行動は反事実的未来を示唆する。 - CtrlWAM は摂動行動をシミュレータで実行し、そのノイズ付き視覚結果とペアにすることでこの不整合を解消。 - これにより、より正確な行動予測、生成動画と行動の一致、供給コマンドへの追従性が向上。 - ロボティクス実験では運動の忠実性と制御性が向上。 - マッチしたコントロールにより、off-path renders がコマンド追従と操作忠実性に寄与することが支持される。

3. 技術・手法の肝は?

- 摂動行動をシミュレータで実行し、そのノイズ付き視覚結果をペアにして WAM を学習。 - 動画と行動の異なるノイズ除去要件に対応するため、warped video-action noise schedules を導入。 - このスケジュールは行動予測の進化に応じて視覚レイアウトの応答性を保つことを目的とする。 - 行動インターフェースを ego-only 制御から可変数のエージェントストリームに拡張。 - これにより単一モデルで複数エージェントの予測または指令された未来を表現可能。 - 具体的なネットワーク構造や損失関数の詳細は要旨からは不明。

4. どうやって有効だと検証した?

- 運転実験とロボティクス実験を実施。 - 運転実験では、より正確な行動予測、生成動画と行動の一致、供給コマンドへの追従性が向上。 - ロボティクス実験では、運動の忠実性と制御性が向上。 - マッチしたコントロールにより、off-path renders がコマンド追従と操作忠実性に寄与することを支持。 - 具体的なデータセット名、評価指標、ベースラインとの比較詳細は要旨からは不明。

5. 議論はある?

- 標準的な WAM 学習における摂動行動とノイズ付き動画の不整合を指摘。 - 低ノイズ領域では、ノイズ付き未来フレームからシーン形状や動的挙動が明確に見えるが、摂動行動は反事実的未来を示唆する。 - CtrlWAM はシミュレータで摂動行動を実行し、そのノイズ付き視覚結果とペアにすることでこの問題に対処。 - warped video-action noise schedules の導入により、行動予測の進化に応じて視覚レイアウトの応答性を保つ。 - 行動インターフェースの拡張により、複数エージェントの予測・指令された未来を統一的に表現可能。 - 限界や今後の課題については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照または比較されている研究は明示されていない。 - 関連手法として World Action Models (WAMs)、video-action noise schedules、ego-only control から可変エージェントストリームへの拡張が挙げられる。 - 同分野の定番として World Models、Action-Conditioned Video Prediction、Diffusion Models for Robotics などが考えられるが、要旨に具体的な論文名はない。 - プロジェクトページ: https://ctrl-wam.github.io/ に詳細がある可能性。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chensheng Peng, Wenhao Ding, Ran Tian, Zewei Zhou, Jef Packer, Maximilian Igl, Peter Karkus, Yan Wang, Masayoshi Tomizuka, Boris Ivanovic, Marco Pavone, Yuxiao Chen

分類: cs.CV

原文アブストラクト

World action models (WAMs) jointly predict actions (intent) and visual future (foresight). Standard training adds noise to recorded actions and video simultaneously, but such training paradigms introduce a mismatch: perturbed actions imply counterfactual future visual, while the noised video remains tied to the GT recording. In low-noise regime, the scene geometry and even the dynamic behavior remain clearly visible from the noisy future frames despite the added noise. We present CtrlWAM, which executes perturbed actions in a simulator and pairs them with their noised visual consequences for joint WAM learning. To accommodate the different denoising requirements of video and actions, we introduce warped video--action noise schedules that aim to keep visual layout responsive as action predictions evolve. We further extend the action interface from ego-only control to a variable number of agent streams, allowing a unified model to represent predicted or commanded futures for multiple agents. Driving experiments show more accurate action forecasts, closer agreement between generated video and actions, and better following of supplied commands; robotics experiments show stronger motion fidelity and controllability. Matched controls support the benefit of off-path renders for command following and manipulation fidelity. Together, these findings contribute to a more controllable world action model. Project page: https://ctrl-wam.github.io/

関連論文

PR本紙発行元 EmplifAI