日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.01742

世界運動モデル:SE(3)軌道の柔軟な系列モデリング

World Motion Models: Flexible Sequence Modeling of SE(3) Trajectories

シェア:XThreadsFacebookLINEはてブBluesky

動的3D世界の要素を剛体SE(3)軌道として表現し、フローマッチングによる系列モデリングで予測・補間・制御・IK・リターゲティング・方策学習を統一的に扱う手法を提案。

詳しい要約

1. どんなもの?

- 動的3D世界の生成priorとしてWorld Motion Models (WMMs)を提案。 - 要素を剛体SE(3)軌跡の集合で近似し、4Dモデリングの最小かつ表現力あるprimitiveとする。 - この表現でarticulated objects、human bodies、hand-object interactions、piecewise-rigid scene dynamics、camera motion、robot states/actionsを単一空間に統合。 - これらの同時分布を柔軟なsequence modeling問題として扱う。

2. 先行研究と比べてどこがすごい?

- 従来の4Dモデリングや軌跡予測と比べ、SE(3)軌跡で多様な実体を統一的に表現。 - 任意の実体・時間ステップ間のany-to-any marginal conditioningを可能に。 - 単一ネットワークでfuture prediction、motion infilling、model-predictive control、inverse kinematics、cross-embodiment retargeting、policy learningをマスク変更のみで実行。 - 6つの3D vision/robotics応用で強い性能を示す。

3. 技術・手法の肝は?

- 動的シーン要素を剛体SE(3)軌跡の集合として近似。 - 同時分布をflow-matching with per-token noise levelsでモデル化。 - 非逐次条件付けのためのcontext token mechanismを導入。 - 任意の実体・時間ステップにマスクを適用することで多様なタスクを同一ネットワークで実現。

4. どうやって有効だと検証した?

- 6つの多様な3D visionとrobotics応用実験で検証。 - 具体的タスクはfuture prediction、motion infilling、model-predictive control、inverse kinematics、cross-embodiment retargeting、policy learning。 - これらの実験でWMMsの汎用性と柔軟性、強い性能を実証。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明記されていない。 - 関連手法としてflow-matching、SE(3) trajectory modeling、sequence modeling、model-predictive control、inverse kinematics、cross-embodiment retargeting、policy learningが挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jiahui Lei, Qianqian Wang, Trevor Darrell, Angjoo Kanazawa

分類: cs.RO, cs.CV

原文アブストラクト

Equipping artificial agents with spatial intelligence requires a comprehensive generative prior over the dynamic 3D world. We propose World Motion Models (WMMs) that capture "what was, is, and will be where across time" via sparse SE(3) pose trajectories. WMMs are built on the observation that elements of dynamic scenes can be well approximated by a set of rigid SE(3) trajectories, a minimal yet expressive primitive for 4D modeling. This representation unifies articulated objects, human bodies, hand-object interactions, piecewise-rigid scene dynamics, camera motion, and even robot states and actions into a single shared space. Given this representation, we cast the joint distribution of these entities as a flexible sequence modeling problem, utilizing flow-matching with per-token noise levels. Coupled with a context token mechanism for non-sequential conditioning, this formulation supports any-to-any marginal conditioning across an arbitrary number of entities and time steps. Tasks such as future prediction, motion infilling, model-predictive control, inverse kinematics, cross-embodiment retargeting, and policy learning all reduce to the application of different masks over the same network. Experiments on 6 diverse applications of 3D vision and robotics demonstrate the versatility and flexibility of WMMs with strong performance.

関連論文

PR本紙発行元 EmplifAI