ヒューマノイドの地平線:並列訓練・動的開始・報酬ゲーティングによる全身ロコマニピュレーションのタスク時間拡張
Humanoid Horizon: Extending Task Horizon in Whole-Body Loco-Manipulation via Parallel Training, Dynamic Starting, and Reward Gating
ヒューマノイドが散らかった室内で複数の物体を連続して運搬・配置する長期的全身ロコマニピュレーションを、並列訓練・動的開始・報酬ゲーティングの3機構で実現する統一ポリシーを提案。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Haozhuo Zhang, Qiang Zhang, Jian Tang, Mingzhe Ni, Michele Caprio, Angelo Cangelosi, Wei Pan
分類: cs.RO, cs.GR
原文アブストラクト
Cluttered indoor environments, where large and heavy objects are scattered across diverse surfaces, require humanoid robots to sequentially navigate, grasp, transport, and accurately place each item at its target location within a single uninterrupted episode. This long-horizon, whole-body loco-manipulation task remains a significant challenge for current methods. Previous approaches often suffer from two main issues: easy-reward bias, where training overemphasizes early transport stages at the expense of later ones, and catastrophic forgetting, where focusing on later stages leads to a decline in earlier-stage performance. In this work, we introduce Humanoid Horizon, a unified policy framework designed to overcome these limitations through three interrelated mechanisms. The Parallel Training Strategy organizes $N$ scenes into $S$ concurrent stage streams governed by a shared policy, ensuring all transport stages receive continuous gradient updates and removing the bottleneck of sequential optimization. The Dynamic Starting Mechanism updates each environment's initial state with terminal states from upstream rollouts, gradually broadening transition coverage and enhancing robustness at stage boundaries. Reward Gating sets the reward to zero for the rest of the episode in later-stage streams when the immediately preceding object is displaced beyond a set threshold, so the shared policy learns not to disturb a just-placed object and earlier placements are preserved throughout the episode. Collectively, these strategies achieve per-stage success rates exceeding 80\% on the two-object LHM-Humanoid benchmark (350 training scenes, 66 held-out scenes). As the number of sequentially transported objects grows beyond two, success declines with the horizon, but the degradation is graceful relative to the sharp drop seen in all baselines.
関連論文
- 接触を決定変数とする:脚式ロコマニピュレーションのための能力トレードオフ接触選択ロコマニピュレーション
- HOTICE: 混雑環境における全身ヒューマノイド物体搬送ロコマニピュレーション
- STRIDER: ヒューマノイドロボットのための歩行を活用した多歩容階層型3Dロコマニピュレーションフレームワークロコマニピュレーション
- 360度カメラのみで実現するモジュラー両手ロコマニピュレーションキャプチャのためのキネマティックインターフェースロコマニピュレーション
- 二足歩行モバイルマニピュレータによる全身協調ロコマニピュレーションの学習ロコマニピュレーション
- Adaptive-MHE:移動ホライズン推定を用いた脚式ロコマニピュレーションのためのサンプリングベース適応MPCロコマニピュレーション