潜在世界モデルにおける長期予測のためのロールアウト復号再構成
Rollout-Decoded Reconstruction for Long-Horizon Prediction in Latent World Models
潜在世界モデルの訓練時に、評価時と同じロールアウトを行い、各潜在変数を復号して再構成誤差をペナルティとして加えることで、長期予測性能を向上させる手法を提案した。
著者: Rishi Shah, Rishav Shrestha
分類: cs.LG
原文アブストラクト
A latent world model trains its decoder on latents anchored to observations, then deploys it on the model's own free-running rollout, hundreds of steps past the last observation. Rollout-Decoded Reconstruction (RDR) closes this gap with a single loss term that free-runs the model during training exactly as evaluation will, decodes every rollout latent, and penalizes reconstruction error against ground truth. The term adds no parameters, costs training-time compute only, and reduces to the standard objective at weight zero, so every comparison in this paper is a one-flag A/B. On the chaotic Kuramoto-Sivashinsky equation, RDR raises valid prediction time (the time to first crossing of normalized error 0.5) from $3.87 \pm 0.23$ to $6.97 \pm 0.42$ time units at an identical 193,568 parameters, a $1.80\times$ improvement confirmed on seeds never used in selection and holding in 10 of 10 preregistered configurations at ratios of 1.71-2.50$\times$. The results come from a single system; a sweep in which the advantage grows with latent width is descriptive, and control experiments on two classic tasks are preliminary.