日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2609.29850

BeyondRetarget: 単眼動画から直接ヒューマノイド動作を学習

BeyondRetarget: Learning Executable Humanoid Motions Directly from Monocular Video

シェア:XThreadsFacebookLINEはてブBluesky

人間の動作表現を介さず、単眼RGB動画から直接ヒューマノイドロボットの動作を生成するエンドツーエンドフレームワークを提案し、接触を考慮した最適化で実行成功率と精度を向上させた。

詳しい要約

1. どんなもの?

- 単眼RGB動画から直接humanoid robotの実行可能な動作を学習するend-to-endフレームワークBeyondRetargetを提案。 - 従来のhuman motion representationを経由せず、視覚観測からrobot-oriented implicit representationを直接学習。 - contact-aware motion optimizationにより時間的一貫性と物理的妥当性を向上。 - シミュレーションと実機humanoid robotで動作精度・ロバスト性・実行成功率・低遅延を実証。

2. 先行研究と比べてどこがすごい?

- 従来はhuman motion representationを構築しmotion retargetingでrobot動作に変換。 - 人間とhumanoidのlocomotion機構やjoint degree-of-freedom構成の差異により、生成動作が実行困難。 - human motion estimationの誤差がretargeting段階に伝播し、joint optimizationでも除去不可。 - BeyondRetargetは明示的human representationを捨て、視覚から直接robot-oriented implicit representationを学習しcross-morphology motion structuresを捉える。

3. 技術・手法の肝は?

- 単眼RGB動画からrobot動作を直接マッピングするend-to-end framework。 - 明示的human representationを排除し、視覚観測からrobot-oriented implicit representationを学習。 - cross-morphology motion structuresを捉える。 - contact-aware motion optimization mechanismを設計し、時間的一貫性と物理的妥当性を改善。

4. どうやって有効だと検証した?

- シミュレーション環境と実humanoid robotで実験。 - 生成robot動作の精度とロバスト性が大幅に向上。 - 実行成功率が高く、遅延が低いことを確認。

5. 議論はある?

- 従来手法の問題点として、人間とhumanoidの形態差による実行困難性と誤差伝播を指摘。 - BeyondRetargetはこれらを克服するが、具体的な限界や議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法としてmotion retargeting、human motion estimation、implicit representation learning、contact-aware optimizationが挙げられる。 - 同分野の定番としてhumanoid motion learning、video-based motion imitationが考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Tianyu Xiong, Yi Lu, Jinrui Wang, Ziqi Liang, Dandan Lei, Xiaoyang Zhou, Xiao-xiao Long, Qiu Shen, Xun Cao

分類: cs.RO, cs.CV

原文アブストラクト

Learning executable motions from human videos offers a scalable solution for humanoid robots to acquire demonstration motions. However, existing pipelines typically first construct an explicit human motion representation and then convert it into robot motions via motion retargeting. Although such methods can effectively leverage large volumes of existing human data for training, the substantial differences between humans and humanoid robots in locomotion mechanisms and joint degree-of-freedom configurations make motions generated by this human-representation-centric approach difficult to execute on robots. Furthermore, errors introduced during human motion estimation inevitably propagate to the retargeting stage and cannot be eliminated via joint optimization. We propose BeyondRetarget, an end-to-end framework that directly maps monocular RGB videos to robot motions. Discarding the explicit human representation, this framework learns robot-oriented implicit representations directly from visual observations, enabling the model to capture cross-morphology motion structures. To generate motions more suitable for robot execution, we further design a contact-aware motion optimization mechanism to improve temporal consistency and physical plausibility. Experiments show that BeyondRetarget significantly improves the accuracy and robustness of generated robot motions, while achieving higher execution success rates and lower latency in both simulation environments and real humanoid robots.

関連論文

PR本紙発行元 EmplifAI