日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.03391

観測動画から直接行動方策を事前学習するNAVA-WAM

Native Action-Prior Learning from Videos for World Action Models

シェア:XThreadsFacebookLINEはてブBluesky

行動ラベルなしの観測動画から未来映像の予測を手がかりに行動方策を直接事前学習し、少数の実演でロボット制御に適応させる手法を提案。

詳しい要約

1. どんなもの?

- 観測のみの動画から行動方策を直接事前学習する手法 NAVA-WAM を提案する研究。 - World action models の一種で、将来の視覚ダイナミクスとロボット行動予測を統合する。 - 行動ラベル付きロボット軌道への依存を減らし、スケーラビリティを高めることを狙う。 - 2段階学習:観測のみ動画での事前学習と、行動ラベル付きデモでの後学習。

2. 先行研究と比べてどこがすごい?

- 既存手法は観測動画を視覚表現の事前学習に使うか、潜在行動を推論してロボットコマンドに接地するかに留まる。 - NAVA-WAM は表現から制御への間接的な転移や別個の潜在行動モデルを避ける。 - 観測のみ動画から行動方策を直接事前学習する native action-prior learning を導入。 - 分布内・分布外の両設定で先行手法を一貫して上回る。

3. 技術・手法の肝は?

- 第1段階:観測のみ動画で、視覚遷移に対する future-video flow-matching 監督を transition-structured joint attention を通じて伝播し、Action-DiT を最適化して行動関連の事前分布を学習。 - 第2段階:行動ラベル付きデモで、joint video-action flow matching により Action-DiT をロボット制御向けに後学習。 - asymmetric attention により視覚枝を反復的な行動 denoising から分離し、行動のみの効率的推論を可能にする。

4. どうやって有効だと検証した?

- 広範な実験で、分布内および分布外設定の両方において先行手法を一貫して上回ることを示す。 - 行動ラベル効率の高さを実証。 - 実ロボットへの汎化が有効であることを示す。

5. 議論はある?

- 観測のみ動画から行動方策を直接事前学習する native action-prior learning が有効なアプローチであると結論。 - 行動ラベル付きロボットデータを超えるスケーラブルな道筋を提供する。 - 具体的な限界や失敗事例、計算コストなどの議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている具体的な先行研究名は不明。 - 関連手法として、観測動画を用いる visual representation pretraining、latent action inference、world action models などが挙げられる。 - 同分野の定番として、robot learning における imitation learning や video pretraining 関連の研究が次に読むべき候補。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhaochong An, Fei Zhang, Menglin Jia, Duncan Frost, Zijian Zhou, Yikai Wang, Xudong Wang, Aditya Patel, Belinda Zeng, Tao Xiang, Serge Belongie, Amir Bar, Sen He

分類: cs.CV, cs.RO

原文アブストラクト

World action models integrate future visual dynamics with robot action prediction, but their scalability remains limited by the need for action-annotated robot trajectories. Observation-only videos contain rich evidence about interaction dynamics, but existing approaches typically use them either to pretrain visual representations that must later be adapted for control, or to infer latent actions that are subsequently grounded to robot commands. We present NAVA-WAM, which introduces native action-prior learning by directly pretraining the action policy from observation-only videos, avoiding indirect representation-to-control transfer or a separate latent-action model. Our training consists of two stages. First, we pretrain on observation-only videos, where future-video flow-matching supervision over visual transitions is propagated through transition-structured joint attention to optimize the Action-DiT and learn action-relevant priors. Second, we use action-labeled demonstrations to post-train the Action-DiT for robot control through joint video--action flow matching, while asymmetric attention decouples the visual branch from iterative action denoising and enables efficient action-only inference. Extensive experiments show that NAVA-WAM consistently outperforms prior approaches under both in-distribution and out-of-distribution settings, while demonstrating strong action-label efficiency and effective real-robot generalization. These results establish native action-prior learning as an effective approach to directly pretrain action policies from observation-only videos, providing a scalable path beyond action-labeled robot data.

関連論文

PR本紙発行元 EmplifAI