日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.09940

Juno: 視覚-言語-行動モデルのための予測潜在表現の活用

Juno: Taming Predictive Latents for Vision-Language-Action Models

シェア:XThreadsFacebookLINEはてブBluesky

行動条件付きJEPAを基盤に、事前学習・方策学習・実機展開の3段階で予測潜在表現を統合し、VLAモデルの成功率を向上させたフレームワーク。

詳しい要約

1. どんなもの?

- Vision-Language-Action (VLA) モデル向けの予測潜在表現を扱う統合フレームワーク Juno を提案。 - 単一の action-conditioned JEPA を中核とし、control-aligned な表現バックボーン、予測 teacher、適応可能な dynamics model の三役を担う。 - pretraining、policy learning、deployment の各段階で予測潜在を活用する。

2. 先行研究と比べてどこがすごい?

- 従来の JEPA は表現空間で masked/future observation を予測するが、embodiment 固有制御との不一致、action learning への干渉、分布シフト下の teacher 誤校正という3つの失敗があった。 - Juno はこれらを統合的に解決し、SimplerEnv で最強ベースライン Qwen3GR00T の平均成功率 60.9% を 68.5% に向上、test-time adaptation で 72.7% に到達。 - 実機でも背景・高さ・物体シフト下で 70%–75% を維持し、ベースポリシーが 0% に崩壊する状況で頑健。

3. 技術・手法の肝は?

- pretraining: embodiment-matched trajectories で action-conditioned JEPA を訓練し、dynamic CLS loss で motion-weighted patch dynamics を compact global state に転移。 - policy learning: 現フレームの JEPA patches を VLA perception に融合し、decoupled reasoning branch と separate transformation parameters で future latent states を蒸留して action generation に利用。 - deployment: 失敗 rollout を含む全観測遷移で world model を適応し、適応済み teacher を凍結、LoRA adapters と trainable action head で verified executions 上にポリシーを再整列。expert corrections や task r…

4. どうやって有効だと検証した?

- SimplerEnv で評価し、平均成功率が Qwen3GR00T の 60.9% から 68.5% へ向上、test-time adaptation で 72.7% に到達。 - 実ロボットで背景・高さ・物体のシフト条件下で 70%–75% の成功率を維持し、ベースポリシーは 0% に崩壊。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- Qwen3GR00T(最強ベースラインとして比較) - JEPA(Joint-embedding predictive architectures) - VLA(Vision-Language-Action)モデル - LoRA adapters - SimplerEnv

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yuchen Zhu, Chenyi Xu, Yulin Zhang, Gang Xu, Wentao Zhu

分類: cs.RO, cs.CV

原文アブストラクト

Joint-embedding predictive architectures (JEPAs) predict masked or future observations in representation space, offering a natural source of predictive latents for vision-language-action (VLA) models. Yet making these latents useful across pretraining, policy learning, and deployment requires addressing three failures: mismatch with embodiment-specific control, interference with action learning, and teacher miscalibration under distribution shifts. We introduce Juno, a unified framework built around one action-conditioned JEPA that serves as a control-aligned representation backbone, a predictive teacher, and an adaptable dynamics model. During pretraining, we train it on embodiment-matched trajectories and use a dynamic CLS loss to transfer motion-weighted patch dynamics to a compact global state. During policy learning, we fuse current-frame JEPA patches into VLA perception and use a decoupled reasoning branch with separate transformation parameters to distill future latent states for action generation. During deployment, we adapt the world model on all observed transitions, including failed rollouts, freeze the adapted teacher, and re-align the policy on verified executions using LoRA adapters and a trainable action head, without expert corrections or task rewards. On SimplerEnv, Juno raises average success from $60.9\%$ to $68.5\%$ over Qwen3GR00T, the strongest baseline, and test-time adaptation further reaches $72.7\%$; on a real robot, it retains $70\%$--$75\%$ success under background, height, and object shifts where the base policy collapses to $0\%$.

関連論文

PR本紙発行元 EmplifAI