日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
全身遠隔操作arXiv:2610.07891

リターゲティングを超えて:学習した原子運動プリミティブによる低遅延で堅牢なヒューマノイド全身遠隔操作

Beyond Retargeting: Low-Latency and Robust Humanoid Whole-Body Teleoperation with Learned Atomic Motion Primitives

シェア:XThreadsFacebookLINEはてブBluesky

人間の動作を直接ロボット関節コマンドに変換するリターゲティング不要のポリシーを提案し、運動プリミティブのコードブックで外れ値や部分入力に頑健に対応、低遅延な全身遠隔操作を実現した。

詳しい要約

1. どんなもの?

- 人型ロボットの全身テレオペレーションを、retargetingなしで実現するポリシーを提案。 - 人間の生の動作を直接ロボットの関節コマンドに単一のforward passでマッピング。 - 学習した全身動作プリミティブのcodebookにより、分布外観測を妥当な動作プロトタイプに射影し、部分入力から全身動作を復元。 - Unitree G1を用いたシミュレーションとハードウェア実験で、VR、光学式mocap、text-to-motion生成、単眼動画入力に対応。

2. 先行研究と比べてどこがすごい?

- 既存システムはオンラインmotion retargetingに依存し、遅延や物理的に不可能な目標を生じる問題があった。 - 提案手法はretargeting-freeでオンライン運動学的適応を排除し、単一forward passで直接関節コマンドを生成。 - 多様でノイズの多い部分的な人間動作観測に対し、codebookにより分布外入力を処理し、部分入力から全身動作を復元可能。 - 遅延とロバスト性の両面でベースラインを上回る。

3. 技術・手法の肝は?

- 人間動作からロボット関節コマンドへの直接マッピングを単一forward passで実現するポリシー。 - 全身動作プリミティブのcodebookを学習し、分布外観測を妥当な動作プロトタイプに射影。 - 部分入力から全身動作を復元する機構を備える。 - オンラインmotion retargetingを不要にし、遅延を低減。

4. どうやって有効だと検証した?

- Unitree G1を用いたシミュレーションとハードウェア実験を実施。 - 入力としてVR、光学式mocap、text-to-motion生成、単眼動画を使用。 - ベースラインと比較し、遅延とロバスト性で優位性を確認。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。関連手法としてmotion retargeting、text-to-motion generation、optical mocap、monocular video-based motion captureが挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xiayan Xu, Jiyu Yu, Xingzhou Chen, Siyi Qian, Zongyu Ma, Lilu Liu, Ling Shi, Haodong Zhang

分類: cs.RO

原文アブストラクト

Humanoid whole-body teleoperation translates human motion into stable robot behavior in real time. Existing systems typically rely on online motion retargeting to bridge human--robot morphological differences, but this process adds latency and can produce physically infeasible targets. Meanwhile, diverse, noisy, and partial human-motion observations often fall outside the training distribution, potentially causing unstable robot behavior. We propose a retargeting-free policy that maps raw human motion directly to robot joint commands in a single forward pass, eliminating online kinematic adaptation. To improve robustness, we learn a codebook of full-body motion primitives that projects out-of-distribution observations onto plausible motion prototypes and recovers full-body motion from partial inputs. Experiments on a Unitree~G1 in simulation and on hardware, using virtual reality, optical mocap, text-to-motion generation, and monocular video inputs, show that our method outperforms baselines in latency and robustness.

関連論文

PR本紙発行元 EmplifAI