日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
計画arXiv:2609.38383

ランダム探索から計画を学習する

Learning to Plan from Random Exploration

シェア:XThreadsFacebookLINEはてブBluesky

ランダム探索の観測ペアのみから時間的関係を学習し、目標までの長距離計画を可能にする手法を提案。

著者: Deqian Kong, Guangyan Sun, Sheng Cheng, Sirui Xie, Bo Pang, Jianwen Xie, Tony Geng, Caiwen Ding, Ying Nian Wu

分類: cs.LG, cs.RO, stat.ML

原文アブストラクト

Random exploration reveals how an environment can be traversed before a goal is specified. Can this experience support long-range planning without policy-improvement training? Our random-walk analysis explains what temporal relations contain: short horizons reveal geodesic geometry in the diffusion limit, while longer horizons reveal connectivity between regions before mixing removes these distinctions. We learn these relations with a conditional energy-based model that estimates temporal log-density ratios through horizon-conditioned embeddings. The model is trained on observation pairs by noise-contrastive estimation, without action or reward labels. The planner queries these learned relations at different horizons as it moves toward the goal. At test time, a separate local dynamics model predicts candidate action outcomes, and the temporal model evaluates their progress toward the goal by selecting or aggregating estimated improvements across horizons. The agent executes one action and replans with both models fixed. Experiments demonstrate long-range maze planning from random exploration using states and images. Learned score fields, embedding probes, and planned routes exhibit properties of a multiscale cognitive map. We further demonstrate egocentric navigation from random exploration and manipulation planning from suboptimal data.

関連論文

PR本紙発行元 EmplifAI