決定論的世界のクローニング:長期世界モデルにおける潜在幾何学の決定的役割
Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models
決定論的な3D世界を高忠実度で模倣する世界モデルを構築するため、長期予測の精度を左右するのは力学モデルではなく潜在表現の幾何構造であることを示し、時間的コントラスト学習による幾何正則化(GRWM)を提案した論文。
著者: Zaishuo Xia, Yukuan Lu, Xinyi Li, Yifan Xu, Yubei Chen
分類: cs.LG, cs.AI, cs.CV
原文アブストラクト
A world model is an internal model that simulates how the world evolves. Given past observations and actions, it predicts the future physical state of both the embodied agent and its environment. Accurate world models are essential for enabling agents to think, plan, and reason effectively in complex, dynamic settings. However, existing world models often focus on random generation of open worlds, but neglect the need for high-fidelity modeling of deterministic scenarios (such as fixed-map mazes and static space robot navigation). In this work, we take a step toward building a truly accurate world model by addressing a fundamental yet open problem: constructing a model that can fully clone a deterministic 3D world. 1) Through diagnostic experiment, we quantitatively demonstrate that high-fidelity cloning is feasible and the primary bottleneck for long-horizon fidelity is the geometric structure of the latent representation, not the dynamics model itself. 2) Building on this insight, we show that applying temporal contrastive learning principle as a geometric regularization can effectively curate a latent space that better reflects the underlying physical state manifold, demonstrating that contrastive constraints can serve as a powerful inductive bias for stable world modeling; we call this approach Geometrically-Regularized World Models (GRWM). At its core is a lightweight geometric regularization module that can be seamlessly integrated into standard autoencoders, reshaping their latent space to provide a stable foundation for effective dynamics modeling. By focusing on representation quality, GRWM offers a simple yet powerful pipeline for improving world model fidelity.