日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ロコマニピュレーションarXiv:2610.08320

ヒューマノイドの地平線:並列訓練・動的開始・報酬ゲーティングによる全身ロコマニピュレーションのタスク時間拡張

Humanoid Horizon: Extending Task Horizon in Whole-Body Loco-Manipulation via Parallel Training, Dynamic Starting, and Reward Gating

シェア:XThreadsFacebookLINEはてブBluesky

ヒューマノイドが散らかった室内で複数の物体を連続して運搬・配置する長期的全身ロコマニピュレーションを、並列訓練・動的開始・報酬ゲーティングの3機構で実現する統一ポリシーを提案。

詳しい要約

1. どんなもの?

本論文は、散らかった屋内環境で人間型ロボットが複数の大きく重い物体を、1エピソード内で順に移動・把持・運搬・配置する long-horizon whole-body loco-manipulation タスクに取り組む。この課題に対し、Humanoid Horizon と呼ぶ unified policy framework を提案する。3つの interrelated mechanisms を組み合わせ、2物体の LHM-Humanoid benchmark で per-stage success rate 80% 超を達成したと報告する。物体数が2を超えると success は horizon とともに低下するが、baselines の急落に比べ graceful な劣化を示す。

2. 先行研究と比べてどこがすごい?

先行手法は主に2つの問題を抱える。1つは easy-reward bias で、訓練が early transport stages を過度に重視し later stages が犠牲になる。もう1つは catastrophic forgetting で、later stages に注力すると earlier-stage performance が低下する。本手法は Parallel Training Strategy、Dynamic Starting Mechanism、Reward Gating を組み合わせ、全 transport stages に継続的な gradient updates を与え、stage boundaries の robustness を高め、just-placed object を乱さないよう shared policy を学習させる点で先行研究と異なる。

3. 技術・手法の肝は?

核となるのは3つの interrelated mechanisms。Parallel Training Strategy は N 個の scenes を S 個の concurrent stage streams に組織し、shared policy で統治することで全 transport stages に連続的な gradient updates を保証し、sequential optimization の bottleneck を除去する。Dynamic Starting Mechanism は各 environment の initial state を upstream rollouts の terminal states で更新し、transition coverage を徐々に広げ stage boundaries での robustness を強化する。Reward Gating は later-stage streams において、直前の object が設定 threshold を超えて変位した場合、残りの episode の reward を zero にし、sha…

4. どうやって有効だと検証した?

LHM-Humanoid benchmark 上で検証している。この benchmark は 350 training scenes と 66 held-out scenes から構成される。2-object 設定で per-stage success rates が 80% を超えることを示した。また、sequentially transported objects の数を2より増やすと success は horizon とともに低下するが、その劣化は all baselines に見られる sharp drop と比較して graceful であると報告している。

5. 議論はある?

本論文は easy-reward bias と catastrophic forgetting という2つの主要問題を明示し、それらを3 mechanisms で克服する枠組みを提示している。2-object 設定では高い per-stage success rate を示す一方、物体数が2を超えると success が horizon とともに低下することを認めており、長い horizon への拡張性が課題として残る。ただし、具体的な failure modes や限界の詳細、計算コスト、実機への転移性については要旨からは不明。

6. 次に読むべき論文は?

要旨で参照・比較されている具体的な先行研究名は明示されていない。関連手法として、long-horizon loco-manipulation における curriculum learning、multi-stage reinforcement learning、catastrophic forgetting 対策、shared policy による並列訓練、reward shaping などが挙げられる。同分野の定番としては、humanoid loco-manipulation の benchmark や whole-body control に関する研究が次に読むべき候補となる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Haozhuo Zhang, Qiang Zhang, Jian Tang, Mingzhe Ni, Michele Caprio, Angelo Cangelosi, Wei Pan

分類: cs.RO, cs.GR

原文アブストラクト

Cluttered indoor environments, where large and heavy objects are scattered across diverse surfaces, require humanoid robots to sequentially navigate, grasp, transport, and accurately place each item at its target location within a single uninterrupted episode. This long-horizon, whole-body loco-manipulation task remains a significant challenge for current methods. Previous approaches often suffer from two main issues: easy-reward bias, where training overemphasizes early transport stages at the expense of later ones, and catastrophic forgetting, where focusing on later stages leads to a decline in earlier-stage performance. In this work, we introduce Humanoid Horizon, a unified policy framework designed to overcome these limitations through three interrelated mechanisms. The Parallel Training Strategy organizes $N$ scenes into $S$ concurrent stage streams governed by a shared policy, ensuring all transport stages receive continuous gradient updates and removing the bottleneck of sequential optimization. The Dynamic Starting Mechanism updates each environment's initial state with terminal states from upstream rollouts, gradually broadening transition coverage and enhancing robustness at stage boundaries. Reward Gating sets the reward to zero for the rest of the episode in later-stage streams when the immediately preceding object is displaced beyond a set threshold, so the shared policy learns not to disturb a just-placed object and earlier placements are preserved throughout the episode. Collectively, these strategies achieve per-stage success rates exceeding 80\% on the two-object LHM-Humanoid benchmark (350 training scenes, 66 held-out scenes). As the number of sequentially transported objects grows beyond two, success declines with the horizon, but the degradation is graceful relative to the sharp drop seen in all baselines.

関連論文

PR本紙発行元 EmplifAI