日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
移動操作arXiv:2609.21467

単一モーションクリップからの距離条件付き物体運搬を学習するヒューマノイド移動操作

Learning Distance-Conditioned Object Transport for Humanoid Loco-Manipulation from a Single Motion Clip

シェア:XThreadsFacebookLINEはてブBluesky

単一の動作クリップから、指定した距離に応じて物体を運ぶヒューマノイドの移動操作を学習する手法を提案。

詳しい要約

1. どんなもの?

- ヒューマノイドの loco-manipulation を単一の retargeted motion clip から学習する手法。 - 固定参照で訓練した方策は実演の最終的な運搬結果しか再現できない問題を扱う。 - 中間変位は passage state として観測されるが、運搬終了は端点でのみ実演されるという termination-versus-passage gap を同定。 - Distance-Conditioned Reference Recomposition (DCRR) を提案し、距離条件付きの運搬を可能にする。

2. 先行研究と比べてどこがすごい?

- 単一 motion clip からの motion tracking による loco-manipulation 再現に対し、固定参照では実演の最終結果のみ再現する制限があった。 - DCRR は実演の終了セグメントを中間運搬状態に再配置し、距離条件付き監督を構築。 - 結果として command-dependent transport を実現し、normalized distance MAE 0.15 を達成(source-only behavior cloning は 0.28)。 - RL fine-tuning により command response と実行頑健性がさらに向上。

3. 技術・手法の肝は?

- DCRR: 実演された termination segment を中間運搬状態に再配置し、参照を再構成。 - 凍結した tracking teacher が再構成参照を closed-loop dynamics 下で再生。 - 保持された trajectory を達成された object placement で relabel し、reference-free policy に distill。 - これにより source motion に符号化された interaction behavior から距離条件付き監督を構築。

4. どうやって有効だと検証した?

- Carry, Kick-Push, Crouch-Push, Drag の4つの interaction mode で評価。 - DCRR-BC は normalized distance MAE 0.15 を達成(source-only behavior cloning は 0.28)。 - RL fine-tuning により training simulator および sim-to-sim transfer で command response と実行頑健性が向上。 - ハードウェア実験で4つの interaction mode すべてにおいて transport-distance modulation を実証。

5. 議論はある?

- termination-versus-passage gap の同定と DCRR による解決を提示。 - RL fine-tuning の有効性を training simulator と sim-to-sim transfer で確認。 - ハードウェア実験で距離変調を実証したが、詳細な限界や失敗事例は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: motion tracking, behavior cloning, RL fine-tuning, sim-to-sim transfer。 - 関連手法: retargeted motion clip からの loco-manipulation 学習、reference-free policy distillation。 - 同分野の定番: humanoid loco-manipulation, motion imitation, reinforcement learning for robotics。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yuhyeon Hwang, Daniel Sungho Jung, YongHyeok Seo, Mingi Jung, Chang Nho Cho, Jung-Hoon Hwang, Dongin Shin

分類: cs.RO

原文アブストラクト

Motion tracking can reproduce humanoid loco-manipulation from a single retargeted motion clip, but a policy trained on a fixed reference primarily reproduces its demonstrated transport outcome. Although the source trajectory visits intermediate object displacements, transport termination is demonstrated only at its endpoint. We identify this mismatch as the termination-versus-passage gap: intermediate displacements are observed as passage states rather than termination-complete outcomes. We introduce Distance-Conditioned Reference Recomposition (DCRR), which relocates the demonstrated termination segment to intermediate transport states. A frozen tracking teacher replays the recomposed references under closed-loop dynamics, and the retained trajectories are relabeled by their achieved object placements and distilled into a reference-free policy. This procedure constructs distance-conditioned supervision from the interaction behavior encoded in the source motion. Across Carry, Kick-Push, Crouch-Push, and Drag, DCRR-BC produces command-dependent transport with an overall normalized distance mean absolute error (MAE) of 0.15, compared with 0.28 for source-only behavior cloning. RL fine-tuning further improves the command response and execution robustness in the training simulator and under sim-to-sim transfer. Finally, hardware experiments demonstrate transport-distance modulation across all four interaction modes.

関連論文

PR本紙発行元 EmplifAI