どれだけ近いかではなく、あとどれだけか:潜在世界モデルの計画のための学習型時間距離指標
How Long, Not How Close: A Learned Temporal Metric for Planning in Latent World Models
潜在世界モデルによる計画で、ゴールまでの残りステップ数を反映する時間距離を学習し、既存モデルを凍結したまま計画コストに組み込むことで、遠いゴールでも成功率を大幅に改善した研究。
著者: Lama Moukheiber, Haotian Xue, Yongxin Chen
分類: cs.LG
原文アブストラクト
Latent world models plan by rolling a frozen predictor forward under candidate action sequences and ranking the candidates by the latent distance between their imagined end state and the goal. However, this ranking breaks down when the goal lies several plans away, because the latent distance measures how closely an end state resembles the goal rather than how far it remains from reaching it. To address this, we propose TEMPO, a temporal-distance planning objective that leaves the world model untouched, learns only from the demonstrations already used to train it, and adds negligible cost to the planner's search. TEMPO learns a small map of the frozen latent in which the distance between two states of an episode reflects the number of environment steps between them, and blends this distance into the planner's cost. It requires no rewards, policies or success labels and, being a cost rather than a model, applies to frozen world models with one latent vector per state that plan by a latent distance. We evaluate TEMPO on eleven simulated environments (e.g., maze navigation, tabletop pushing, robotic arm control and three-dimensional manipulation) with the LeWM and PLDM planners. With a small MLP that adds at most 0.3% to a plan's arithmetic, TEMPO improves both planners at every goal distance, including the one-plan setting of their evaluations, raises LeWM from 36% to 99% on TwoRoom three plans from the goal, and remains competitive on a broad range of 2D and 3D navigation, reaching and manipulation tasks.
関連論文
- 遠くへ届くには近くを狙え:凍結世界モデルは思うより賢く計画できるモデルベース計画
- プリズマティック・ワールドモデル:ハイブリッド系の計画のための合成可能なダイナミクス学習モデルベース計画
- ロバスト計画のための因果構造分布の学習モデルベース計画