日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2608.18948

RoboEdit: 人間の操作動画をスケーラブルなロボット経験へ変換する

RoboEdit: Turning Human Manipulation Videos into Scalable Robot Experience

シェア:XThreadsFacebookLINEはてブBluesky

人間の操作動画を、ロボットの学習に使えるアクション整合性のある物理的に妥当なロボット動画へ変換するビデオ編集スイートを提案。大規模データセットと編集エンジンを構築し、実世界の操作タスクで有効性を実証した。

詳しい要約

1. どんなもの?

RoboEditは、人間の操作ビデオをロボットのビデオに変換する編集スイートである。RGBビデオから3Dインタラクションを再構成・リターゲットする自動パイプラインRoboEdit-ADCを導入し、174Kの整列したビデオペア(14Mフレーム)からなる大規模データセットRoboEdit-14Mを生成する。コアの編集エンジンRoboEdit-Transは、クロスエンボディメント適応モジュールと3Dロボット状態デコーダを統合し、時間的一貫性を保ちながら外観と動作を適応させる。

2. 先行研究と比べてどこがすごい?

従来のロボットデータ収集はコストが高く、特定のエンボディメントに依存するが、人間のビデオは豊富に存在するもののロボット訓練には直接使えなかった。RoboEditは、人間のビデオを物理的に妥当でアクション整合的なロボットビデオに変換することで、未ラベルの人間ビデオをスケーラブルなロボット学習の監督信号として活用できる点が革新的である。

3. 技術・手法の肝は?

手法の肝は、RoboEdit-ADCによる3Dインタラクションの再構成とリターゲティング、およびRoboEdit-Transによるクロスエンボディメント適応である。RoboEdit-Transは、時間的一貫性を保つための適応モジュールと、フレームごとの手の状態を復元する3Dロボット状態デコーダを備え、構造化されたモーション監督を提供する。

4. どうやって有効だと検証した?

実験では、編集品質の最先端性を達成し、実世界の操作タスクにおける下流のロボット制御ポリシーをサポートすることを示した。具体的な評価指標や比較対象は要旨からは不明だが、編集品質と実世界タスクでの有効性が検証されている。

5. 議論はある?

要旨からは、データセットの規模や多様性、編集品質の優位性が示されているが、限界や議論については明記されていない。例えば、生成されたビデオの物理的妥当性の検証方法や、異なるエンボディメント間の一般化の限界などは不明である。

6. 次に読むべき論文は?

要旨で参照されている関連研究は明示されていないが、同分野の定番として、人間の操作ビデオからのロボット学習(例えば、Learning from Human Videos)や、ビデオ編集・生成(例えば、Video Diffusion Models)に関する論文が考えられる。具体的には、RoboEditが基盤とする可能性のある3D再構成やリターゲティングの手法(例えば、Hand pose estimation, Motion retargeting)を読むと良い。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yaowei Guo, Zeng Tao, Yuxin Jiang, Yunuo Chen, Zhiyang Dou, Yuxiang Ma, Yin Yang, Demetri Terzopoulos, Ying Jiang, Chenfanfu Jiang

分類: cs.RO

原文アブストラクト

Collecting robot hand-object interaction data is costly and embodiment-specific, yet abundant human-object videos remain unusable for robot training. We present RoboEdit, a human-to-robot video editing suite that transforms human manipulation videos into action-consistent, physically plausible robot videos with aligned 3D hand states. To enable scalable supervision, we introduce RoboEdit-ADC, an automatic pipeline that reconstructs and retargets 3D interactions from RGB videos across embodiments. This pipeline generates RoboEdit-14M, a large-scale dataset of 174K aligned video pairs (14M frames) spanning seven robot embodiments, diverse scenes, and interaction types. The core editing engine, RoboEdit-Trans, employs cross-embodiment adaptation modules to preserve temporal coherence while adapting appearance and motion. It further integrates a 3D Robot-State Decoder to recover per-frame hand states for structured motion supervision. Experiments show that RoboEdit achieves state-of-the-art editing quality and supports downstream robot control policies in real-world manipulation tasks. Ultimately, the RoboEdit suite unlocks the vast potential of unlabeled human videos, providing scalable, high-fidelity visual and 3D motion supervision for generalizable robot learning.

関連論文