日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.35761

DexRoam: 一人称全身人間デモンストレーションからの移動型両腕巧緻マニピュレーション学習

DexRoam: Learning Mobile Bimanual Dexterous Manipulation from Egocentric Whole-Body Human Demonstrations

シェア:XThreadsFacebookLINEはてブBluesky

VRヘッドセットと頭部搭載ステレオカメラだけで全身の人間デモを収集し、身体・動作意味・時間の3段階アライメントでロボット動作空間に変換することで、移動しながらの両腕巧緻操作をVLAポリシーで学習可能にした。

詳しい要約

1. どんなもの?

- モバイル双腕巧緻マニピュレーションを、egocentricな全身人間デモンストレーションから学習するシステムDexRoamを提案。 - 移動・全身運動・指先巧緻性を単一軌道内で連続的に協調させるタスクを対象とする。 - 人間の全身運動を連続かつ結合したままhuman-to-robot転移する完全なシステム。 - 消費者向けVRヘッドセットと頭部搭載ステレオカメラのみで、tracker-freeに全身操作デモを収集。

2. 先行研究と比べてどこがすごい?

- 従来のegocentric人間デモ活用は、人間運動を単純化して転移を容易にし、タスクが依存する微細で結合した構造を捨てていた。 - DexRoamは全身運動を連続かつ結合したまま転移し、微細な全身運動構造を保持する点が異なる。 - 外部カメラやmotion trackerを必要とせず、VRヘッドセットと頭部ステレオカメラのみで収集可能。 - 人間とロボットのデモを標準VLAポリシーでjointlyに学習できる。

3. 技術・手法の肝は?

- tracker-free capture system: 消費者向けVRヘッドセットと頭部搭載ステレオカメラのみで全身操作デモを収集。 - 3つの明示的alignment段階: embodiment、action-semantic、temporalを実施。 - 取得運動をロボットaction spaceへ写像し、微細な全身運動を保持。 - 人間とロボットのデモを標準VLAポリシーでjointlyに学習可能にする。

4. どうやって有効だと検証した?

- 異なるVLA backboneを用いた実世界実験を実施。 - 人間デモが学習パラダイム間で一貫してポリシー学習を改善。 - GR00T N1.7で平均成功率29%→56%、pi0.5で32%→57%に向上。 - ロボットデモ半量でrobot-only訓練と同等を達成。 - Ablationで各alignment段階が必須であることを確認。

5. 議論はある?

- 微細な運動構造を保持した人間デモが、スケーラブルな全身モバイルマニピュレーションに有効である可能性を示す。 - 各alignment段階の必要性がablationで支持される。 - 限界や失敗事例、計算コスト、安全性などの詳細は要旨からは不明。

6. 次に読むべき論文は?

- GR00T N1.7 - pi0.5 - 標準VLAポリシー - egocentric human demonstrationを用いたhuman-to-robot transferの関連研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Rui Zhou, Yibo Yuan, Junkai Zhao, Fangyuan Zhao, Xiaoguang Zhao, Shanghang Zhang, Sirui Han

分類: cs.RO

原文アブストラクト

Mobile bimanual dexterous manipulation requires continuous coordination of locomotion, whole-body motion, and finger-level dexterity within a single trajectory, creating a severe robot demonstration bottleneck. Egocentric human demonstrations offer a scalable alternative, but prior approaches ease the transfer by simplifying human motion, discarding exactly the fine-grained, coupled structure such tasks depend on. We present DexRoam, a complete system for learning mobile bimanual dexterous manipulation from human demonstrations, in which whole-body motion remains continuous and coupled throughout the human-to-robot transfer process. To enable scalable collection of whole-body human manipulation demonstrations, we develop a tracker-free capture system using only a consumer VR headset and a head-mounted stereo camera, without external cameras or motion trackers. We then perform three explicit alignment stages---embodiment, action-semantic, and temporal---to map captured motion into the robot action space, preserving fine-grained whole-body motion and allowing human and robot demonstrations to be jointly learned by standard VLA policies. Real-world experiments with different VLA backbones show that human demonstrations consistently improve policy learning across training paradigms, raising average success from 29% to 56% on GR00T N1.7 and from 32% to 57% on pi0.5, while matching robot-only training with half the robot demonstrations. Ablations confirm that each alignment stage is necessary. These results highlight the potential of human demonstrations for scalable whole-body mobile manipulation with preserved fine-grained motion structure.

関連論文

PR本紙発行元 EmplifAI