潜在計画のための時間的直線化
Temporal Straightening for Latent Planning
世界モデルを用いた潜在計画のための表現学習を改善する手法を提案。人間の視覚処理における知覚的直線化仮説に着想を得て、曲率正則化により潜在軌道を局所的に直線化し、計画の安定性と成功率を向上させる。
著者: Ying Wang, Oumayma Bounou, Gaoyue Zhou, Randall Balestriero, Tim G. J. Rudner, Yann LeCun, Mengye Ren
分類: cs.LG
原文アブストラクト
Learning good representations is essential for latent planning with world models. While pretrained visual encoders produce strong semantic visual features, they are not tailored to planning and contain information irrelevant -- or even detrimental -- to planning. Inspired by the perceptual straightening hypothesis in human visual processing, we introduce temporal straightening to improve representation learning for latent planning. Using a curvature regularizer that encourages locally straightened latent trajectories, we jointly learn an encoder and a predictor of a Joint-Embedding Predictive Architecture (JEPA) world model. We show that reducing curvature this way makes the Euclidean distance in latent space a better proxy for the geodesic distance and improves the conditioning of the planning objective. We demonstrate empirically that temporal straightening makes gradient-based planning more stable and yields significantly higher success rates across a suite of goal-reaching tasks. Our code is available at https://agenticlearning.ai/temporal-straightening.
関連論文
- 世界モデルのためのテキスト信念状態:厳密な媒介下での識別可能な表現学習世界モデル/表現学習
- 潜在世界の探査:潜在表現における創発的離散記号と物理構造世界モデル/表現学習