日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.18232

UMI-Bridge: 人間とロボットの操作データを動作で結ぶ潜在表現アライメント

UMI-Bridge: Action-Anchored Latent Alignment across Human and Robot Manipulation Data

シェア:XThreadsFacebookLINEはてブBluesky

UMIを中間ドメインとして使い、見た目ではなく動作の等価性に基づいて人間とロボットの操作データの潜在表現を揃え、VLAポリシーの学習を改善する手法を提案。

詳しい要約

1. どんなもの?

- 実ロボットのデモが限られる中、ロボット無しで収集した人間の操作データ(egocentric videos や handheld UMI デモ)を活用する研究。 - UMI を中間ドメインとして、視覚的外観ではなく操作運動の等価性に基づき表現を整列する UMI-Bridge を提案。 - 人間操作データで dual-view latent action model (LAM) を訓練し、その wrist teacher と dynamics model を凍結して VLA の post-training を正則化。 - 実ロボット3タスクで平均成功率 91.7% を達成。

2. 先行研究と比べてどこがすごい?

- 従来は視点・embodiment・action supervision の違いから、外観ではなく操作運動に基づく表現整列が困難だった。 - UMI-Bridge は UMI を中間ドメインに据え、action equivalence に基づく整列を実現。 - 同一 UMI・ロボットデータでの Naive Co-training の 73.3% に対し 91.7% と大幅改善。 - ロボットデモ 25% + UMI データで full-data Robot-only baseline を上回るデータ効率を実現。

3. 技術・手法の肝は?

- UMI action supervision が潜在表現を end-effector motion と gripper behavior にアンカー。 - 同期した head-wrist 観測と paired ego-UMI clips が視点・ドメイン間の整列を支援。 - ロボットデモ無しの人間操作データで dual-view LAM を訓練。 - wrist teacher と dynamics model を凍結し、UMI・ロボットデータでの VLA post-training を正則化。 - 共有 wrist interface により推論アーキテクチャを変えずに訓練時監督を両ドメインに適用。

4. どうやって有効だと検証した?

- 実ロボット3タスクで評価し、平均成功率 91.7%(Naive Co-training は 73.3%)。 - データ効率2タスクで、ロボットデモ 25% + UMI データが full-data Robot-only baseline を上回る。 - タスク特化ロボットデモ無しの UMI デモから学習した追加2タスクで 85% と 90% の成功率。

5. 議論はある?

- 結果は data-efficient robot learning と UMI-to-robot task transfer に対する action-anchored latent alignment の有効性を支持。 - 限界や失敗事例、計算コスト、一般化範囲についての議論は要旨からは不明。

6. 次に読むべき論文は?

- Naive Co-training(比較ベースライン) - Robot-only baseline(比較ベースライン) - Universal Manipulation Interface (UMI) - latent action model (LAM) - vision-language-action (VLA) post-training

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Haiyi Liu, Jingming Ma, Ke Rui, Yuteng Wei, Yuan Ma, Yushen Zuo, Honglong Tian, Haoran Jia, Weitao Zhou, Jiawei Wang, Minglei Li, Shiyi Chen, Haiyan Mao, Jiaqi Zhang, Chun Zhang

分類: cs.RO

原文アブストラクト

Real-robot demonstrations are limited, motivating the use of human manipulation data collected without robots, including egocentric videos and handheld Universal Manipulation Interface (UMI) demonstrations. However, differences in viewpoint, embodiment, and available action supervision make it difficult to align representations across these sources according to manipulation motion rather than visual appearance. We introduce UMI-Bridge, which uses UMI as an intermediate domain to align representations according to action equivalence rather than pixel similarity. UMI action supervision anchors the latent representation to end-effector motion and gripper behavior, while synchronized head-wrist observations and paired ego-UMI clips support alignment across views and domains. We train a dual-view latent action model (LAM) on human manipulation data without robot demonstrations, then freeze its wrist teacher and dynamics model to regularize vision-language-action (VLA) post-training on UMI and robot data. The shared wrist interface enables this training-time supervision across both domains while preserving the policy's standard inference architecture. Across three real-robot tasks, UMI-Bridge achieves 91.7% mean success versus 73.3% for Naive Co-training with matched UMI and robot data. On two data-efficiency tasks, it surpasses a full-data Robot-only baseline using 25% of the robot demonstrations together with UMI data. It also achieves 85% and 90% success on two additional tasks learned from UMI demonstrations without task-specific robot demonstrations. These results support action-anchored latent alignment for data-efficient robot learning and UMI-to-robot task transfer.

関連論文

PR本紙発行元 EmplifAI