日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
強化学習/報酬設計arXiv:2610.04250

ROOT: ユーザー指定の身体動作のための報酬発見フレームワーク

ROOT: Discovering Rewards for User-Specified Embodied Behaviors

シェア:XThreadsFacebookLINEはてブBluesky

動画言語モデルと実験木探索を組み合わせ、自然な歩容や姿勢など視覚的に認識しやすい身体動作を報酬関数として自動発見する手法を提案。シミュレーションと実機の四足歩行ロボットで有効性を実証した。

詳しい要約

1. どんなもの?

- 強化学習における報酬設計の難しさを解決するため、ユーザー指定の身体動作を実現する報酬関数を発見するフレームワーク「ROOT」を提案。 - 自然言語記述から報酬関数を合成するLLMベース手法では、歩容や姿勢などの微妙な行動特性を捉えきれない問題に対処。 - 観測に基づく探索を通じて、学習方策をユーザー意図に沿わせる報酬関数を発見する。

2. 先行研究と比べてどこがすごい?

- 既存のLLMベース報酬生成手法と比較して、ユーザー意図により適合する行動を生成。 - 移動完全性精度で最大86.8%を達成し、Vid-LLMによる行動整合性スコアを3.56/5から4.14/5へ16.5%改善。 - 人間評価でもペア比較の51-63%でROOTが好まれる。

3. 技術・手法の肝は?

- 報酬設計を観測ガイド付き探索として定式化し、永続的な実験ツリーを構築。 - 実験ツリーには報酬プログラム、学習済み方策、ロールアウト観測、およびビデオ言語モデルが抽出した行動的洞察を保存。 - スカラー訓練統計のみに頼らず、行動的失敗を診断し、その後の報酬改善を導く。

4. どうやって有効だと検証した?

- 4つの身体(シミュレーションのHopper, HalfCheetah, Ant, Unitree Go2、および実世界のUnitree Go2)にわたる7タスクで評価。 - 移動完全性精度とVid-LLM行動整合性スコアで既存LLMベース手法を上回る。 - 人間によるペア比較評価でもROOTが好まれる割合が51-63%であることを確認。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されているLLMベース報酬生成手法(例:Eureka, Language to Rewardsなど)や、ビデオ言語モデル(Vid-LLM)を活用した報酬設計に関する研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Eren Sadikoglu, Aditya Taparia, Xinyuan Liu, Ransalu Senanayake

分類: cs.RO

原文アブストラクト

Reinforcement learning for embodied control remains constrained by the difficulty of reward specification. Although recent large language model (LLM)-based methods can synthesize reward functions from natural-language descriptions, they often fail to capture subtle behavioral properties that humans care about, such as natural gait, posture, and movement style. This limitation arises because many desired behaviors are easier to recognize visually than to encode in a reward function. We introduce Reward Optimization via Observable Trees (ROOT), a framework for discovering reward functions that align learned policies with user-specified embodied behaviors. Rather than relying solely on scalar training statistics, ROOT casts reward design as an observation-guided search over a persistent experiment tree that stores reward programs, trained policies, and rollout observations, together with behavioral insights distilled by a video-language model, to diagnose behavioral failures and guide subsequent reward refinements. We evaluate ROOT on seven tasks across four embodiments: simulated Hopper, HalfCheetah, Ant, Unitree Go2, and as well as the real-world Unitree Go2. ROOT produces behaviors that better align with user intent than those generated by existing LLM-based reward-generation methods, achieving up to 86.8% locomotion-completeness accuracy and improving Vid-LLM behavioral alignment from 3.56/5 to 4.14/5, a 16.5% improvement over baselines. Human evaluations further support these results, with ROOT preferred in 51-63% of pairwise comparisons.

関連論文

PR本紙発行元 EmplifAI