日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.30594

HuGo: ヒューマノイドの移動操作のための全身ポリシーコードを設計するLLM

HuGo: LLMs as Whole-Body Policy Code Designers for Humanoid Loco-Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

LLMがタスク記述から高レベルポリシーコードを生成し、ロールアウトで改良することで、報酬設計や実演なしにヒューマノイドの移動操作を実現する階層的手法を提案。

詳しい要約

1. どんなもの?

ヒューマノイドのloco-manipulationを対象に、LLMで高レベルpolicyコードを生成する階層的手法HuGoの提案。 - タスク記述から実行可能なclosed-loop高レベルpolicyコードをLLMが生成。 - 凍結した低レベルwhole-body policyの上に構築。 - タスク毎のreward設計やdemonstration収集を不要にする狙い。

2. 先行研究と比べてどこがすごい?

従来はreward engineeringやdemonstrationに続くタスク特化学習が必要で、新タスクへの拡張コストが高い。 - HuGoはタスク毎のreward設計やdemonstration収集なしで性能を達成。 - 5つのsimulationタスクで高レベルRL baselineを大幅に上回り、demonstration-based baselineに迫る。 - 同一refinement loopを実世界rolloutに適用し、専門家demonstrationやpolicy再学習なしで転移性能を改善。

3. 技術・手法の肝は?

階層的アプローチ。 - LLMがタスク記述・observation・command仕様から高レベルpolicyコードを生成。 - 凍結された低レベルwhole-body policy上でclosed-loop実行。 - rolloutの数値trajectoryと選択video frameを用いてfeedbackとtargeted code updateを行いpolicyをrefine。

4. どうやって有効だと検証した?

5つのsimulationタスクで2種類の低レベルpolicyを用いて評価。 - 高レベルRL baselineを大幅に上回る。 - demonstration-based baselineに迫る性能。 - simulation生成policyのhardwareへのzero-shot転移を実証。 - 実世界rolloutへの同refinement loop適用で転移性能がさらに改善。

5. 議論はある?

タスク毎のreward設計やdemonstration収集なしで性能を達成できる点を主張。 - 実世界rolloutへのrefinement適用が専門家demonstrationやpolicy再学習なしで有効。 - 限界や失敗事例、計算コスト、安全性などの議論は要旨からは不明。

6. 次に読むべき論文は?

要旨で参照/比較されている研究: 高レベルreinforcement learning baseline、demonstration-based baseline。 - 関連手法: reward engineering、demonstration、task-specific training、frozen low-level whole-body policy、LLMによるpolicy code generation。 - 同分野の定番: humanoid loco-manipulation、hierarchical RL、whole-body control。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Seoyeon Choi, Shizhao Ye, Nicholas Bui, Aayushi Shrivastava, Kanghyun Ryu, Dhruva Tirumala, Markus Wulfmeier, Negar Mehr

分類: cs.RO

原文アブストラクト

For humanoids to be useful in everyday environments, they must perform a wide range of tasks that couple locomotion and manipulation. Existing approaches commonly acquire a loco-manipulation policy through reward engineering or demonstrations followed by task-specific training, making it costly to scale to new tasks. In this work, we propose a hierarchical approach to humanoid loco-manipulation that eliminates these per-task requirements. HuGo, Humanoid policy code Generation, uses a Large Language Model (LLM) to generate executable, closed-loop high-level policy code from a task description on top of a frozen low-level whole-body policy. Given the task, observation, and command specifications, the LLM constructs the task logic in code. HuGo then refines the policy from its rollouts using numerical trajectories and selected video frames to produce feedback and targeted code updates. Across five simulation tasks, using two different low-level policies, HuGo substantially outperforms a high-level reinforcement learning baseline and approaches the performance of a demonstration-based baseline. We achieve this level of performance without task-specific reward design or demonstration collection. We further demonstrate zero-shot transfer of simulation-generated policies to hardware and show that applying the same refinement loop to real-world rollouts can further improve transfer performance without expert demonstrations or policy retraining. Project website is https://iconlab.negarmehr.com/HuGo/

関連論文

PR本紙発行元 EmplifAI