日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデルarXiv:2610.12299

細粒度身体インタラクションを伴うマルチエージェント自己中心的世界モデル

Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction

シェア:XThreadsFacebookLINEはてブBluesky

複数エージェントが共有環境で細かい身体動作を通じて相互作用する様子を、各エージェント視点の映像として同時生成する世界モデルME-Worldを提案し、共有世界の一貫性や行動制御の精度を向上させた。

詳しい要約

1. どんなもの?

- 複数エージェントが共有環境で細粒度の身体動作を通じて相互作用する際の、一人称視点(egocentric)観測を予測する世界モデル。 - 従来の単一エージェント中心のegocentric world modelを拡張し、複数エージェントの同期されたego-stream生成として定式化。 - 提案手法ME-Worldは、共有トークン列で複数のego streamを同時にdenoiseし、全エージェントのtarget-view poseで条件付け、共有環境メモリで生成を根拠付ける。

2. 先行研究と比べてどこがすごい?

- 既存のmulti-agent world modelはlocomotion、camera control、離散コマンドなどの粗い行動に依存し、細粒度のembodied interactionは未探索だった。 - ME-Worldは細粒度行動による相互作用を扱い、cross-view action consistency、shared-environment consistency、interaction-induced state updatesの一貫した伝播を実現。 - 実験でshared-world consistency、action control、identity preservation、video qualityが既存手法より向上。

3. 技術・手法の肝は?

- 複数のego streamを共有トークン列でjointly denoiseするアーキテクチャ。 - 各streamを全エージェントのtarget-view poseで条件付け、cross-view action consistencyを確保。 - 共有環境メモリで生成をgroundingし、shared-environment consistencyとinteraction-induced state updatesの一貫した伝播を支援。

4. どうやって有効だと検証した?

- 実データと合成のmulti-agentデータで訓練・評価。 - 環境、更新、identityの一貫性を測るshared-world consistency metricsを新たに導入。 - 実験により、shared-world consistency、action control、identity preservation、video qualityの改善を確認。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。関連手法としてmulti-agent world model、egocentric world model、embodied interaction、video diffusion modelなどが挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Dahyun Chung, Siyoon Jin, Hyunwook Choi, Honggyu An, Junyoung Seo, Hyunsung Kim, Seung Wook Kim, Seungryong Kim

分類: cs.CV, cs.AI

原文アブストラクト

Egocentric world models predict first-person observations conditioned on an agent's actions, but most focus on a single agent. Real embodied settings often involve multiple agents that act and interact within a shared environment. Existing multi-agent world models rely on coarse actions like locomotion, camera control, or discrete commands, leaving fine-grained embodied interactions underexplored. We formulate multi-agent egocentric world modeling as synchronized ego-stream generation for multiple agents interacting through fine-grained actions in a shared world. This requires cross-view action consistency, shared-environment consistency, and consistent propagation of interaction-induced state updates. We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory. We train and evaluate on real and synthetic multi-agent data and introduce shared-world consistency metrics for environment, update, and identity consistency. Experiments show ME-World improves shared-world consistency, action control, identity preservation, and video quality over existing methods.

関連論文

PR本紙発行元 EmplifAI