日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
目標条件付き方策学習arXiv:2610.09247

目標条件付き方策学習におけるホライズンの情報的呪い

An Informational Curse of Horizon in Goal-Conditioned Policy Learning

シェア:XThreadsFacebookLINEはてブBluesky

目標再ラベリングのホライズンが長くなると方策の汎化性能が低下する現象を「情報的呪い」として特定し、入力ヤコビアンの蒸留で改善できることを示した。

詳しい要約

1. どんなもの?

- ゴール条件付き政策学習における新たな「horizonの情報的呪い」を特定 - ゴールrelabeling horizonを長くすると政策の汎化と性能が大幅に低下 - 訓練時のhorizonとテスト時のhorizonを分離した制御実験で検証 - BCとRLの比較、条件付き相互情報量の減少、入力Jacobianの感度シフトを分析 - 短horizon政策のJacobianを蒸留することで性能改善

2. 先行研究と比べてどこがすごい?

- 従来のhorizonの呪いはTD backupのバイアス蓄積やadvantage推定のノイズとして説明 - 本研究は追加の「情報的呪い」を特定し、政策汎化の低下を明示 - ゴールrelabeling horizonが一般ist政策学習の重要因子であることを示唆 - BCとRLの性能差をhorizon依存の情報量減少で説明 - 入力Jacobianの蒸留という新たな改善手法を提案

3. 技術・手法の肝は?

- oracle plannerを用いた制御実験で訓練とテストのゴールhorizonを分離 - ゴール条件付きBCとRL政策を比較 - 行動とhindsight-relabeled goals間の条件付き相互情報量を分析 - 政策の入力Jacobianを測定し、ゴールと状態への感度シフトを評価 - 短horizon政策のJacobianを長horizon政策に蒸留する手法を適用

4. どうやって有効だと検証した?

- 一連の制御実験でhorizon依存の性能劣化を確認 - 近傍サブゴールでの評価でもBCが深刻な劣化を示すことを検証 - RL目的が劣化を緩和することを実証 - 条件付き相互情報量のhorizon依存減少を経験的に確認 - 入力Jacobianの蒸留が特にcombinatorial manipulationタスクで性能向上をもたらすことを検証

5. 議論はある?

- ゴールrelabeling horizonがオフラインデータからの一般ist政策学習の重要考慮点 - BCとRLの性能差の原因を情報理論的に説明 - 入力Jacobianの感度シフトが性能劣化に関与 - 蒸留による改善はhorizonの呪いへの対処法として有望 - 要旨からは不明な点として、他のタスクやデータセットへの一般化可能性

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: 明示的な参照はないが、goal-conditioned behavioral cloning (BC) と reinforcement learning (RL) の関連研究 - 関連手法: hindsight relabeling, temporal-difference backups, advantage estimation - 同分野の定番: goal-conditioned policy learning, offline reinforcement learning, combinatorial manipulation

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: John L. Zhou, Yuxuan Dong, Jonathan C. Kao

分類: cs.LG, cs.AI, cs.RO

原文アブストラクト

The difficulty of learning goal-reaching policies is often attributed to a "curse of horizon" that manifests as bias accumulation in temporal-difference backups and noisy advantage estimates. In this work, we identify an additional informational curse of horizon in goal-conditioned policy learning, where increasing the goal relabeling horizon can significantly reduce policy generalization and performance. Through a series of controlled experiments with oracle planners, we decouple the goal horizons sampled during training from those that the policy is asked to reach at test time. Even when evaluated only on a sequence of nearby subgoals, goal-conditioned behavioral cloning (BC) policies suffer from severe, training horizon-dependent performance degradation that is mitigated by reinforcement learning (RL) objectives. We explain this phenomenon as a horizon-dependent decrease in the conditional mutual information between actions and hindsight-relabeled goals, and find empirically that both BC and RL policies trained on longer-horizon goals exhibit a shift in sensitivity from goal to state information, as measured by the policy's input Jacobians. Motivated by this observation, we find that distilling the input Jacobians of short-horizon policies into long-horizon policies yields significant performance gains, especially in combinatorial manipulation tasks. Taken together, our results highlight goal relabeling horizon as an important consideration when learning generalist policies from offline data.

PR本紙発行元 EmplifAI