日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
arXiv:2608.03387

RoboReact: Agentic Skill Distillation from Generated Egocentric Videos for Generalizable Whole-Body Manipulation

RoboReact: Agentic Skill Distillation from Generated Egocentric Videos for Generalizable Whole-Body Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

詳しい要約

1. どんなもの?

RoboReactは、単一のegocentric RGB-D観測から全身ヒューマノイド操作スキルを自動合成するフレームワーク。ビデオ生成モデルで人間の操作動画を生成し、depth-aware 3D reconstructionで幾何学を保つinteraction keyframesを抽出、高自由度のヒューマノイドプラットフォームにリターゲットする。オンラインのobject-centric re-groundingとvision-language modelによるrefinement loopで物理実行のギャップを埋め、whole-body controllerで実行する。テレオペレーションや人間のデモを必要とせず、多様な物体配置に一般化し、外乱から回復できる。

2. 先行研究と比べてどこがすごい?

従来のヒューマノイドスキル獲得は高価なハードウェアデータ収集とアノテーションに依存していた。ビデオ生成モデルを活用した先行研究は、想像上の行動を実行可能な全身スキルに転送する点が未開拓だった。RoboReactは、生成動画から幾何学を保つキーフレームを抽出し、オンライン再接地とVLMガイドのrefinementで物理的実行可能性を高め、テレオペやデモなしで実機で一般化と外乱回復を実証した点が新しい。

3. 技術・手法の肝は?

手法の肝は、(1) 単一のegocentric RGB-D観測から人間の操作動画を生成し、depth-aware 3D reconstructionで手と物体の相互作用幾何学を保つキーフレームを抽出すること。(2) 高DoFヒューマノイドへのリターゲティング。(3) オンラインのobject-centric re-groundingで幾何学的ミスマッチを補正。(4) vision-language modelによるrefinement loopで実行偏差を適応。(5) whole-body controllerで全身協調操作を実現。

4. どうやって有効だと検証した?

実ヒューマノイドロボットでの実験により、多様な物体構成への一般化と、実行中の外乱からの回復を実証した。テレオペレーションや人間のデモを必要としないことを確認。具体的な評価指標や比較対象は要旨からは不明。

5. 議論はある?

要旨からは、生成動画の品質やリターゲティングの精度、実機での成功率などの定量的評価が不明。また、単一のRGB-D観測に依存するため、複雑な環境や動的物体への適用限界が考えられるが、要旨では議論されていない。

6. 次に読むべき論文は?

要旨で参照されている関連研究として、ビデオ生成モデル(例: video generative models)、vision-language models、whole-body control、skill retargeting、egocentric video understandingなどが挙げられる。具体的な論文名は不明だが、これらの分野の定番論文を読むとよい。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Shuliang He, Shuai Wang, Bo Yue, Junchi Teng, Changyu Wang, Guiliang Liu

分類: cs.RO

原文アブストラクト

Humanoid robots have the potential to perform dexterous manipulation in human environments, yet acquiring diverse and generalizable skills remains costly due to expensive hardware data collection and labor-intensive annotation. Recent advances in video generative models provide a promising opportunity to synthesize rich manipulation experiences from visual observations, but transferring such imagined behaviors into executable whole-body humanoid skills remains largely unexplored. In this work, we present RoboReact, a framework that automatically synthesizes whole-body humanoid manipulation skills from a single egocentric RGB-D observation. RoboReact generates human manipulation videos, extracts geometry-preserving interaction keyframes through depth-aware 3D reconstruction, and retargets them to high-DoF humanoid platforms while preserving hand-object interaction geometry. To bridge the gap between imagined plans and physical execution, RoboReact performs online object-centric re-grounding and leverages a vision-language model-guided refinement loop to adapt skills under geometric mismatch and execution deviations. The refined skills are executed through a whole-body controller, enabling coordinated whole-body manipulation and dexterous interaction. Experiments on real humanoid robots demonstrate that RoboReact generalizes across diverse object configurations and robustly recovers from execution disturbances without requiring teleoperation or human demonstrations. These results highlight the potential of combining generative models, vision-language reasoning, and closed-loop control for scalable humanoid skill acquisition.