日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
移動操作arXiv:2609.38172

反実仮想動画生成によるスケーラブルなヒューマノイドの移動操作

Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

少数の実動画から反実仮想の人間-物体インタラクション動画を大量生成し、実→シミュ→実のパイプラインで物理的に妥当な軌道に変換して、多様な物体の運搬を実機で実現する方策を学習した。

詳しい要約

1. どんなもの?

- 本論文は、ヒューマノイドのloco-manipulationスキルを視覚模倣で教えるためのフレームワークPRISMを提案する。 - 実世界の少数の動画を増幅し、多様な訓練データを生成するreal-to-sim-to-realアプローチ。 - 具体的には、video-to-video (V2V) 生成により数百のcounterfactualなhuman-object interaction動画を生成。 - その後、contact-anchored real-to-simパイプラインで人間と物体の動きを再構築し、物理的に妥当な軌道にリターゲット。 - このデータで単一のポリシーを訓練し、実ロボットに実世界ファインチューニングなしで展開。 - オンボード深度のみを用いて、箱、バレル、ビン、ボールなどの物体を新しいインスタンス、サイズ、初期配置で持ち上げ、運び、落とすことを実現。

2. 先行研究と比べてどこがすごい?

- 従来の視覚模倣では、多様で高品質なインタラクション動画(全身が明確で物体とのインタラクションが隠れていないクリップ)の収集がスケーリングの障壁であった。 - PRISMは、少数の実動画からcounterfactual動画を生成することでこの制限を克服し、大規模で多様な訓練セットを構築。 - これにより、実世界でのファインチューニングなしで、カテゴリ内の未見物体に汎化する単一ポリシーを訓練可能にした点が先行研究と比べて優れている。 - 具体的な比較対象は要旨からは不明。

3. 技術・手法の肝は?

- 技術の肝は、real-to-sim-to-realフレームワークPRISMにある。 - まず、少数の実例動画からvideo-to-video (V2V) 生成を用いて数百のcounterfactualなhuman-object interaction動画を生成。 - 次に、contact-anchored real-to-simパイプラインで人間と物体の動きを再構築し、不完全なビデオデータを物理的に妥当な軌道にリターゲット。 - このcounterfactual動画間のintra-class variabilityを利用して、単一のポリシーを訓練。 - ポリシーはオンボード深度観測のみを使用。

4. どうやって有効だと検証した?

- 実ロボットにポリシーを展開し、実世界でのファインチューニングなしで評価。 - オンボード深度のみを用いて、箱、バレル、ビン、ボールなどの物体を、新しいインスタンス、サイズ、初期配置で持ち上げ、運び、落とすタスクを実行。 - これにより、パイプライン全体の有効性を実証。

5. 議論はある?

- 要旨からは、議論や限界についての明示的な記述は不明。 - ただし、実世界ファインチューニングなしで汎化性能を示した点が強調されている。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、visual imitation learning、video-to-video generation、real-to-sim-to-real transfer、humanoid loco-manipulationなどが挙げられる。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zihan Wang, Zhen Wu, Pieter Abbeel, Rocky Duan, Jitendra Malik, Carmelo Sferrazza, C. Karen Liu, Guanya Shi, Angjoo Kanazawa

分類: cs.RO, cs.CV, cs.GR

原文アブストラクト

Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. However, collecting diverse, high-quality interaction videos, such as clips that clearly show a person's full body and unoccluded interactions with objects, poses a practical barrier to scaling this approach. We propose PRISM, a real-to-sim-to-real framework that overcomes this limitation by amplifying a handful of real videos into a large, diverse training set. PRISM first generates hundreds of diverse "counterfactual" human-object interaction videos via video-to-video (V2V) generation from a few exemplar real videos. Our contact-anchored real-to-sim pipeline then reconstructs both human and object motions, retargeting this imperfect video data into physically plausible trajectories. The intra-class variability across these counterfactual videos lets us train a single policy that generalizes to unseen objects within each category. We demonstrate the full pipeline by deploying this policy on a real robot without any real-world fine-tuning. Using only onboard depth observations, our humanoid picks up, carries, and drops objects, including boxes, barrels, bins, and balls, across novel instances, sizes, and initial configurations.

関連論文

PR本紙発行元 EmplifAI