日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2608.18787v1

Dream2Reward: 正のデモンストレーションからの遷移整合報酬モデルによるロボット操作

Dream2Reward: Transition-Alignment Reward Models from Positive Demonstrations for Robotic Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

成功デモから遷移レベルの潜在変位を予測し、観測された動作との整合性で密な報酬を生成する手法を提案。失敗アノテーション不要で報酬ハッキングを軽減し、実ロボット操作を含む下流性能を向上させる。

著者: Haoyu Zhang, Zecui Zeng, Bin Wang, Lusong Li, Liang Lin, Long Cheng

分類: cs.RO

原文アブストラクト

Learning robotic policies requires dense rewards that remain informative when behavior departs from successful demonstrations. Progress-based rewards estimate how far an observation has advanced along a nominal successful trajectory, but may remain high after an incorrect transition. We introduce Dream2Reward, which learns a language-conditioned successful latent transition field from positive demonstrations. Given the visual history up to a transition start, the model predicts the latent displacement associated with successful execution and scores the observed displacement through signed directional and symmetric magnitude agreement. This transition-level comparison penalizes wrong-direction, overshooting, and stagnant motion even when the resulting observation appears to show progress. Dream2Reward requires no failure annotations, progress labels, or synthetic negatives, and produces a dense causal reward. Across mechanism diagnostics and shared-trajectory evaluations, it provides stronger success-failure separation and more informative feedback on low-quality behavior than progress-based alternatives. Across online and offline policy learning, the same frozen reward model reduces reward hacking and supports stronger downstream performance, including in real-robot manipulation. These results show that comparing realized motion with predicted successful change provides an effective way to convert positive demonstrations into dense rewards for robot learning.

関連論文