日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
模倣学習arXiv:2610.10857

自己教師ありキーフレーム発見による長期依存行動クローニング

Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning

シェア:XThreadsFacebookLINEはてブBluesky

ランダムにサンプリングした過去観測から情報量の多いキーフレームを自己教師ありで発見し、それを条件に行動クローニングを行うことで、長期依存タスクでも性能劣化なく汎化できる手法を提案。

著者: Prabin Kumar Rath, Omkar Patil, Nakul Gopalan

分類: cs.AI, cs.RO

原文アブストラクト

Behavior cloning (BC) in non-Markovian environments is a challenging problem because policies have to reason over contextual information over long horizons. Existing policy architectures rely on recurrent or attention-based mechanisms to capture long-term dependencies. However, recurrent models suffer from hidden-state collapse and gradient instability under backpropagation through time, while attention-based models are fundamentally limited by context length. To address these issues, we propose Keyframe Mnemonics, a novel self-supervised method that $\textit{discovers}$ a set of information-critical observations ($\textit{mnemonics}$) by learning an objective from randomly sampled past observations and using it as a reward for keyframe selection. We then train a BC policy that conditions on the discovered keyframes to model the action distribution. Under certain task-structure assumptions, our formulation provides context retention guarantees over an infinite horizon, while maintaining a small set of decision-relevant keyframes in the policy's working memory. We evaluate our method on synthetic memory domains, where mnemonic-conditioned BC policies achieve $100$% success rates (SR) and generalize to horizons orders of magnitude beyond training without performance degradation. Additionally, we evaluate on memory-intensive robot manipulation benchmark, achieving a $13.9$% average absolute SR improvement over the strongest baseline across $23$ tasks and retaining $80$% SR at $20\times$ longer horizons on a real robot. Code and videos are available at https://keyframe-mnemonics.github.io.

関連論文

PR本紙発行元 EmplifAI