日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
模倣学習arXiv:2601.00452

軌道レベルの生成埋め込みによる観察からの模倣学習

Imitation from Observations with Trajectory-Level Generative Embeddings

シェア:XThreadsFacebookLINEはてブBluesky

拡散モデルで軌道データを潜在空間に埋め込み、専門家の状態密度を推定することで、不完全なオフラインデータからでも密で滑らかな報酬を構築する観察ベース模倣学習手法を提案。

著者: Yongtao Qu, Shangzhe Li, Weitong Zhang

分類: cs.LG

原文アブストラクト

We consider the offline imitation learning from observations (LfO) where the expert demonstrations are scarce and the available offline suboptimal data are far from the expert behavior. Many existing distribution-matching approaches struggle in this regime because they impose strict support constraints and rely on brittle one-step models, making it hard to extract useful signal from imperfect data. To tackle this challenge, we propose TGE, a trajectory-level generative embedding for offline LfO that constructs a dense, smooth surrogate reward by estimating expert state density in the latent space of a temporal diffusion model trained on offline trajectory data. By leveraging the smooth geometry of the learned diffusion embedding, TGE captures long-horizon temporal dynamics and effectively bridges the gap between disjoint supports, ensuring a robust learning signal even when offline data is distributionally distinct from the expert. Empirically, the proposed approach consistently matches or outperforms prior offline LfO methods across a range of D4RL locomotion and manipulation benchmarks.

関連論文