日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
移動操作arXiv:2610.02196

InterEvolve: ヒューマノイド移動操作のための報酬プログラムのテスト時進化

InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

訓練済みコントローラを再訓練せずに、LLMと数値最適化で報酬プログラムをテスト時に進化させ、ヒューマノイドの移動操作タスクを解く手法を提案。

詳しい要約

1. どんなもの?

- 本研究は、humanoid loco-manipulation における test-time evolution を扱う。 - コントローラが未訓練のタスクを、既存スキルを再利用し、自己の試行から改善し、学習内容を保持することで、再訓練なしに解決する。 - 鍵となるのは、planning と control の間のインターフェースで、接触の多い多段階の相互作用を指定できる表現力を持ち、実行可能かつ測定可能で、実行フィードバックが経験から planning を導く。 - InterEvolve はこのインターフェースを2つの要素で実現する。 - 物体認識型の forward-backward (FB) behavioral foundation model を開発。 - タスクを reward program として指定。 - 実験では、人間設計の報酬が FB モデルの loco-manipulation 能力の多くを未活用であることを示し、InterEvolve が進化させたプログラムがそれを解放する。 - 多様なタスク、複雑なシーン、長期的な構成の行動をシミュレーションで…

2. 先行研究と比べてどこがすごい?

- 先行研究と比べて、test-time evolution により、再訓練なしに新しいタスクを解決する点がすごい。 - 人間設計の報酬では FB モデルの能力が十分に活用されないのに対し、InterEvolve はそれを解放し、時には新規戦略を生み出す。 - 従来の手法では、コントローラの既存スキルを再利用して新しいタスクに適応することが難しかったが、本研究は planning と control のインターフェースを通じてそれを可能にする。 - また、LLM エージェントと数値最適化を組み合わせて reward program を進化させる点が新しい。

3. 技術・手法の肝は?

- 技術の肝は、物体認識型の forward-backward (FB) behavioral foundation model と reward program の組み合わせ。 - FB モデルは、凍結された body prior 上の object residuals を用いて、身体や物体に関する新しい報酬を test time で loco-manipulation 行動に変換する。 - タスクは reward program として指定:段階的な報酬、完了条件、調整可能な定数を含む。 - LLM エージェントが実行フィードバックと検証済みプログラムのスキルライブラリを参照してプログラム構造を文脈内で修正し、数値オプティマイザが定数を調整する。 - 各候補は並列シミュレーションシナリオで検証され、プログラムはコントローラの既存運動能力を誘導、再利用、構成する新しい方法を探索し、反復的に改善する。

4. どうやって有効だと検証した?

- 実験により、人間設計の報酬が FB モデルの loco-manipulation 能力の多くを未活用であることを示した。 - InterEvolve が進化させたプログラムがそれを解放し、時には新規戦略を通じて実現することを確認。 - 多様なタスク、複雑なシーン、長期的な構成の行動をシミュレーションで生成。 - 進化したスキルが物理的な Unitree G1 上で自己中心的なオンボード知覚から自律実行されることを検証。

5. 議論はある?

- 議論の詳細は要旨からは不明。 - ただし、人間設計の報酬の限界と、InterEvolve による能力解放の可能性が示唆されている。 - 新規戦略の生成や、シミュレーションから実機への転移が議論の焦点となる可能性がある。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、forward-backward (FB) behavioral foundation model、reward program、LLM エージェント、数値最適化、humanoid loco-manipulation の test-time adaptation などが挙げられる。 - 同分野の定番として、強化学習ベースの loco-manipulation、sim-to-real 転移、行動基礎モデル (behavioral foundation models) などが次に読むべき論文として考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhuo Lin, Sirui Xu, Liuyu Bian, Yu-Xiong Wang, Liang-Yan Gui

分類: cs.RO, cs.CV, cs.GR

原文アブストラクト

We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns, without retraining. Our key insight is that a broad controller already holds much of the competence a new task needs, and that this competence becomes accessible through an interface between planning and control that is expressive enough to specify contact-rich, multi-stage interactions, yet executable and measurable enough that execution feedback can guide planning from experience. InterEvolve realizes this interface with two components. First, we develop an object-aware forward-backward (FB) behavioral foundation model, whose object residuals on a frozen body prior turn a new reward about the body or objects into loco-manipulation behavior at test time. Second, we specify tasks as reward programs: staged rewards with completion conditions and tunable constants. A large language model (LLM) agent revises the program structure in context, drawing on execution feedback and a skill library of verified programs, while a numerical optimizer tunes its constants. With every candidate verified across parallel simulation scenarios, the program explores new ways to induce, repurpose, and compose the controller's existing motor competence for the task at hand, and thus improves over iterations. Experiments show that human-designed rewards leave much of the FB model's loco-manipulation competence untapped, whereas the programs InterEvolve evolves release it, sometimes through novel strategies. It further produces behaviors for diverse tasks, complex scenes, and long-horizon compositions in simulation, and evolved skills run autonomously on a physical Unitree G1 from egocentric onboard perception.

関連論文

PR本紙発行元 EmplifAI