日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.37250

V-JEPAポリシー:予測的視覚潜在表現上に構築する効果的な世界行動モデル

V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents

シェア:XThreadsFacebookLINEはてブBluesky

凍結したV-JEPA 2.1エンコーダの潜在空間上に世界行動モデルを構築し、将来潜在予測器とフローマッチング行動エキスパートを下流で一括学習することで、LIBERO等で競争力のある性能を達成した研究。

詳しい要約

1. どんなもの?

- V-JEPA Policy は、World-Action Model (WAM) を構築するためのシンプルなフレームワーク。 - 凍結した V-JEPA 2.1 encoder の潜在空間上に WAM を構築する。 - instruction-conditioned future-latent predictor と flow-matching action expert を、単一の下流段階でゼロから同時学習する。 - 総パラメータ 0.9B、うち学習可能 0.6B。 - LIBERO、LIBERO-Plus、RoboCasa-GR1 で代表的な WAM や vision-language-action ベースラインと競争力のある性能を示す。

2. 先行研究と比べてどこがすごい?

- 従来の WAM は大規模事前学習済みの video generator や image-editing model を適応し、予測知識とその学習に用いたモデルを継承していた。 - 本研究は、完全な事前学習済み視覚生成モデルを継承せず、大規模予測事前学習で得られた predictive visual latent space が WAM 学習の十分な基盤になるかを問う。 - 同じ下流フレームワークと学習予算で視覚基盤を比較し、V-JEPA latents が discriminative、reconstructive、video-understanding-oriented の代替より効果的、特に分布シフト下で有効と同定。 - 行動ラベルなしの DROID video-instruction pairs で predictor を事前学習し WAM に適応すると、下流制御と out-of-distribution 汎化が大幅に向上。

3. 技術・手法の肝は?

- 凍結した V-JEPA 2.1 encoder の潜在空間を基盤とする。 - instruction-conditioned future-latent predictor と flow-matching action expert を単一の下流段階でゼロから同時学習。 - predictor の future-informed context key-value states が action generation を条件付ける。 - 総 0.9B パラメータ、学習可能 0.6B。 - 行動ラベルなしの DROID video-instruction pairs で predictor を事前学習し、WAM に適応可能。

4. どうやって有効だと検証した?

- LIBERO、LIBERO-Plus、RoboCasa-GR1 で代表的な WAM および vision-language-action ベースラインと比較。 - 同じ下流フレームワークと学習予算で視覚基盤を比較し、V-JEPA latents の有効性を検証。 - 分布シフト下での性能を評価。 - DROID video-instruction pairs での predictor 事前学習と WAM 適応による下流制御と out-of-distribution 汎化の向上を検証。

5. 議論はある?

- 予測視覚潜在空間が、タスク固有デモンストレーションからの効果的な WAM 学習の基盤となることを示す。 - 広範な in-the-wild 動画から獲得した future-modeling knowledge の転移可能性を示す。 - 完全な事前学習済み視覚生成モデルを継承せずに WAM を学習できる可能性を提示。 - 具体的な限界や議論の詳細は要旨からは不明。

6. 次に読むべき論文は?

- V-JEPA 2.1 - LIBERO - LIBERO-Plus - RoboCasa-GR1 - DROID - vision-language-action ベースライン - video generator や image-editing model を基盤とする WAM

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yang Zhang, Jiangyuan Zhao, Chenyou Fan, Jiayu Hu, Xiu Yuan, Chenjia Bai, Xiu Li

分類: cs.CV, cs.AI, cs.RO

原文アブストラクト

World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrained at scale, a prominent line of recent WAMs inherits both predictive knowledge and the models in which it was learned. We ask whether a predictive visual latent space induced by large-scale predictive pretraining can instead provide a sufficient foundation for effective WAM learning without inheriting a complete pretrained visual generative model. To answer this question, we introduce V-JEPA Policy, a simple framework that builds a WAM on the latent space of a frozen V-JEPA 2.1 encoder. An instruction-conditioned future-latent predictor and a flow-matching action expert are jointly learned from scratch in a single downstream stage, with the predictor's future-informed context key--value states conditioning action generation. With 0.9B total parameters, of which 0.6B are trainable, V-JEPA Policy achieves competitive performance with representative WAM and vision-language-action baselines across LIBERO, LIBERO-Plus, and RoboCasa-GR1. Comparing visual foundations under the same downstream framework and training budget identifies V-JEPA latents as more effective than the discriminative, reconstructive, and video-understanding-oriented alternatives, particularly under distribution shifts. Beyond task-specific learning, pretraining the predictor on DROID video--instruction pairs without action labels and adapting it into a WAM yields substantial gains in downstream control and out-of-distribution generalization. Together, these findings establish predictive visual latents as a foundation for effective WAM learning from task-specific demonstrations and for transferring future-modeling knowledge acquired from broader in-the-wild videos. Our code is available at https://github.com/breez3young/VJEPA-Policy.

関連論文

PR本紙発行元 EmplifAI