日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデル/操作arXiv:2608.29242v1

AnyWorld: 因子分解された自己中心的世界モデルによるクロス身体汎化

AnyWorld: Factorized Egocentric World Models for Cross-Embodiment Generalization

シェア:XThreadsFacebookLINEはてブBluesky

人間の単一インタラクション動画を、ロボット固有の多様なロールアウトに拡張する世界モデルフレームワークを提案。行動・カメラ・身体の因子に分解し、身体・視点・シーンの再構成を可能にする。

詳しい要約

1. どんなもの?

AnyWorldは、単一の人間のインタラクションを多様なロボット固有のロールアウトに拡張する、クロスエンボディメント世界モデルフレームワークである。アクション、カメラ、エンボディメントの3要素にインタラクションを分解し、それぞれを独立に再構成することで、ペアの人間-ロボットデモなしに、ロボットドメインのデータを生成する。大規模な人間のインタラクションプリトレーニングと混合エンボディメントのファインチューニングで訓練される。

2. 先行研究と比べてどこがすごい?

従来のロボット学習は、データ量と多様性の不足が課題であり、特にエンボディメント、視点、シーンの多様性が欠けていた。人間のegocentricビデオは豊富な物理的インタラクションを提供するが、単一の身体、カメラ軌道、環境しか捉えられない。AnyWorldは、この狭い経験を分解し再構成することで、ペアデータなしにロボット固有のデータを生成できる点が新しい。

3. 技術・手法の肝は?

手法の核は、インタラクションをアクション制御、カメラ制御、エンボディメントコンテキストの3要素に分解することである。アクション制御は動作構造を捉え、カメラ制御は視点の進化を指定し、エンボディメントコンテキストは動作する身体とそのインタラクション幾何学を定義する。この分解により、エンボディメント、視点、シーンの要素を独立に再構成でき、単一モデルで多様なロボットドメインのデータを生成できる。

4. どうやって有効だと検証した?

RoboCasa GR1テーブルトップベンチマークと実機IRONヒューマノイドロボットで、生成データが操作性能を向上させることを実証した。さらに、ペアのない人間の経験をロボット固有のビデオ-アクションペアに再構成し、ポリシーのギャップをターゲットにできるかテストした。制御されたIRON介入により、誤った完了事前分布を修正し、言語に基づく空間ターゲット選択を確立した。アクションのみの反事実介入では後者を確実に学習できず、アクションのキャリブレーションと視覚的再構成の両方が必要であることを示した。

5. 議論はある?

要旨からは、分解の一般性や他のエンボディメントへの適用範囲、生成データの品質の限界、実世界でのスケーラビリティなどについての議論は不明である。また、アクションと視覚の再構成の必要性は示されたが、その相互作用の詳細や、他のタスクでの有効性については言及されていない。

6. 次に読むべき論文は?

要旨で参照されている関連研究は明示されていないが、同分野の定番として、世界モデル(World Models)、クロスエンボディメント学習(Cross-Embodiment Learning)、人間のビデオからのロボット学習(Learning from Human Videos)、拡散モデルによるビデオ生成(Diffusion Models for Video Generation)などが挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Cheng Chen, Jerry Bai, Jiacheng Wei, Boyu Chen, Xiaoji Zheng, Fan Wu, Minghao Yang, Tianrun Chen, Ruibo Li, Xiaoyu Yue, Xiaoyang Guo, Yixiao Ge, Guosheng Lin, Fayao Liu

分類: cs.RO

原文アブストラクト

Collecting contact-rich robot experiences at scale remains a major bottleneck for generalizable manipulation. Beyond data quantity, robot learning also requires diverse experiences across embodiments, viewpoints, and scenes. Human egocentric videos provide abundant physical interactions, but each video captures only a narrow slice of experience under a single body, camera trajectory, and environment. We propose AnyWorld, a cross-embodiment world modeling framework that expands a single human interaction into diverse robot-native rollouts without paired human-robot demonstrations. Our model factorizes an interaction into action, camera, and embodiment: action controls capture the motion structure, camera controls specify viewpoint evolution, and the target embodiment context defines the acting body and its interaction geometry. This formulation enables independent recomposition of embodiment, viewpoint, and scene factors, allowing a single model to generate many robot-domain experiences while preserving the underlying dynamics and object interactions. We train the model with large-scale human interaction pretraining followed by mixed-embodiment fine-tuning. Experiments show that our model supports controllable recomposition across embodiments, viewpoints, and scenes, and we further demonstrate that the generated data can improve manipulation performance on the RoboCasa GR1 tabletop benchmark and a real IRON humanoid robot. Beyond aggregate gains, we test whether unpaired human experience can be recomposed into robot-native video-action pairs that target a policy gap. Controlled IRON interventions correct a spurious completion prior and establish language-grounded spatial target selection; an action-only counterfactual intervention fails to learn the latter reliably, showing that both action calibration and visual recomposition are necessary.

関連論文