ROOT: ユーザー指定の身体動作のための報酬発見フレームワーク
ROOT: Discovering Rewards for User-Specified Embodied Behaviors
動画言語モデルと実験木探索を組み合わせ、自然な歩容や姿勢など視覚的に認識しやすい身体動作を報酬関数として自動発見する手法を提案。シミュレーションと実機の四足歩行ロボットで有効性を実証した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Eren Sadikoglu, Aditya Taparia, Xinyuan Liu, Ransalu Senanayake
分類: cs.RO
原文アブストラクト
Reinforcement learning for embodied control remains constrained by the difficulty of reward specification. Although recent large language model (LLM)-based methods can synthesize reward functions from natural-language descriptions, they often fail to capture subtle behavioral properties that humans care about, such as natural gait, posture, and movement style. This limitation arises because many desired behaviors are easier to recognize visually than to encode in a reward function. We introduce Reward Optimization via Observable Trees (ROOT), a framework for discovering reward functions that align learned policies with user-specified embodied behaviors. Rather than relying solely on scalar training statistics, ROOT casts reward design as an observation-guided search over a persistent experiment tree that stores reward programs, trained policies, and rollout observations, together with behavioral insights distilled by a video-language model, to diagnose behavioral failures and guide subsequent reward refinements. We evaluate ROOT on seven tasks across four embodiments: simulated Hopper, HalfCheetah, Ant, Unitree Go2, and as well as the real-world Unitree Go2. ROOT produces behaviors that better align with user intent than those generated by existing LLM-based reward-generation methods, achieving up to 86.8% locomotion-completeness accuracy and improving Vid-LLM behavioral alignment from 3.56/5 to 4.14/5, a 16.5% improvement over baselines. Human evaluations further support these results, with ROOT preferred in 51-63% of pairwise comparisons.
関連論文
- 報酬設計エージェント:強化学習のための報酬設計強化学習/報酬設計
- 強化学習のための段階遷移型高密度報酬モデリング強化学習/報酬設計
- 生成的エピソードガイダンスによる二重粒度コントラスト報酬で効率的な身体性RL強化学習/報酬設計
- TimeRewarder:受動的動画からフレーム間時間距離で密報酬を学習強化学習/報酬設計
- グラフ思考による報酬進化:強化学習のための二段階言語モデルフレームワーク強化学習/報酬設計
- ORSO: オンライン報酬選択と方策最適化による報酬設計の加速強化学習/報酬設計