日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.24033

Imagine-RL:予測残差信頼度で導くクロスアテンションによる世界モデル拡張VLA強化学習

Imagine-RL: Residual-Confidence-Guided Cross-Attention for World-Model-Augmented VLA Reinforcement Learning

シェア:XThreadsFacebookLINEはてブBluesky

凍結したVLA方針をノイズ空間強化学習で微調整する際、視覚・力覚の潜在世界モデルで将来の結果を予測し、その予測残差を信頼度として活用することで接触の多い操作タスクの成功率を大幅に向上させた。

詳しい要約

1. どんなもの?

- 接触を伴うマニピュレーションのための強化学習手法 - 凍結したVision-Language-Action (VLA) ポリシーをノイズ空間で微調整 - 行動条件付きの視覚・トルク想像(visual-torque imagination)を導入 - 未来の視覚・接触結果を考慮して行動評価を改善 - 実ロボット4タスクで成功率を大幅向上

2. 先行研究と比べてどこがすごい?

- 既存のノイズ空間RL(DSRLなど)は批評家が未来の結果を無視 - 本研究は未来の視覚・トルク表現を想像し批評家に統合 - DSRL比で平均成功率+23.6%、VLAベースライン比+60% - 100 RL trajectoriesのみで効率的に学習

3. 技術・手法の肝は?

- 凍結したvisual-torque latent world model (VTLWM) を利用 - 各候補行動チャンクに対し、自己回帰的に未来のコンパクト表現を予測(ピクセル再構成なし) - 現在の画像・状態・行動クエリが観測履歴と予測未来にクロスアテンション - 前ウィンドウの予測残差をトークンごとの信頼度事前分布として利用し、信頼できない未来トークンを抑制 - 現在の証拠と予測結果を組み合わせて批評家が行動評価、俳優を監督 - VLAとVTLWMは凍結

4. どうやって有効だと検証した?

- 実ロボット4タスクで各50回の評価試行 - 100 RL trajectoriesのみを使用 - 平均成功率をDSRL比+23.6%、VLAベースライン比+60%改善 - 具体的なタスク内容や評価指標の詳細は要旨からは不明

5. 議論はある?

- 要旨からは不明 - 限界や失敗事例、計算コスト、一般化性に関する議論は記述なし

6. 次に読むべき論文は?

- DSRL (Noise-space reinforcement learning for VLA) - VLA (Vision-Language-Action) ベースライン - Visual-torque latent world model (VTLWM) 関連研究 - その他のworld-model-augmented RLやcontact-rich manipulationの研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Kejia Hu, Wentong Zhai, Bo Zhao, Shuai Liang

分類: cs.RO

原文アブストラクト

Reliable action evaluation in contact-rich manipulation requires looking beyond the current observation to future visual and contact consequences. Existing noise-space reinforcement learning efficiently steers a frozen Vision-Language-Action (VLA) policy, but its critics largely ignore these consequences. We present Imagine-RL, which augments noise-space VLA post-training with action-conditioned visual-torque imagination. For each candidate action chunk, a frozen visual-torque latent world model (VTLWM) autoregressively predicts compact future representations without pixel reconstruction. A current image-state-action query attends to observed histories and predicted futures, while previous-window prediction residuals provide token-wise confidence priors that suppress unreliable future tokens. By combining current evidence with predicted consequences, the action critic better evaluates candidate actions and supervises the actor, while the VLA and VTLWM remain frozen. Across four real-robot tasks with 50 evaluation trials per task, Imagine-RL uses only 100 RL trajectories and improves the average success rate by (23.6%) over DSRL and by (60%) over VLA baselines.

関連論文

PR本紙発行元 EmplifAI