日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
群制御arXiv:2610.09438

世界モデル計画による制御可能な群衆生成

Controllable Crowd Generation through World-Model Planning

シェア:XThreadsFacebookLINEはてブBluesky

群衆シミュレーションに世界モデルを導入し、再学習なしで実行時に制御目標を変更できるマルチエージェント手法Ctrl-CWMを提案した。

詳しい要約

1. どんなもの?

- ロボットナビゲーション、自動運転、都市計画で重要な群集シミュレーション。 - 既存手法は事前定義された制御設定に依存し、ユーザー指定の新たな目的への柔軟性が低い。 - 本研究では、群集生成と実行時制御を統合したマルチエージェント Controllable Crowd World Model (Ctrl-CWM) を提案。 - 世界モデル原理の「想像された未来を用いた計画」を群集シミュレーションに適応。 - エンコーダ、アクター、クリティック、プランナーから構成。

2. 先行研究と比べてどこがすごい?

- 既存手法は事前定義された制御設定に依存し、新たなユーザー目的への適応が困難。 - Ctrl-CWM は実行時にユーザーコストを導入することで、再トレーニングなしに新しい制御目的を追加可能。 - 最先端手法と比較して、多くの群集リアリズムと衝突メトリクスで優れる。 - シミュレーション中に導入されたユーザー指定目的に群集行動を適応させる。

3. 技術・手法の肝は?

- 世界モデル原理に基づき、想像された未来を用いた計画を群集シミュレーションに適用。 - エンコーダ:人間の運動ダイナミクスの表現を学習。 - アクター:歩行者の変位を提案。 - クリティック:想像された群集軌道を評価。 - プランナー:行動を選択。 - 実世界の歩行者ビデオでの軌道予測により運動ダイナミクスを学習し、エンコーダを凍結して保持。 - アクターは繰り返し状態更新を通じて想像群集軌道を生成。 - プランナーはクリティックのスコアとユーザーコストを組み合わせて行動選択。 - 繰り返し計画でシミュレーション群集を進め、追加のユーザーコストで新制御目的を導入。

4. どうやって有効だと検証した?

- 多様なエージェント到着条件での群集生成と、回避・誘引シナリオでの実行時制御を広範囲に評価。 - 最先端手法と比較し、ほとんどの群集リアリズムと衝突メトリクスで優れる。 - シミュレーション中に導入されたユーザー指定目的に群集行動を適応させることを確認。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究や関連手法は明示されていない。 - 同分野の定番として、Social Force Model、Social GAN、Trajectron++ などが挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: JunGyu Lee, Jisu Shin, Seunghyun Shin, Hae-Gon Jeon

分類: cs.CV

原文アブストラクト

Crowd simulation plays a central role in robot navigation, autonomous driving, and urban planning. For these applications, realistic simulation requires crowds to adapt their behavior to environmental changes and user objectives. However, existing methods that rely on predefined control settings have limited flexibility in accommodating new user-specified objectives. To address this limitation, we propose Ctrl-CWM, a multi-agent Controllable Crowd World Model that integrates crowd generation and run-time control. Our key idea is to adapt the world-model principle of planning using imagined futures to crowd simulation. To this end, Ctrl-CWM consists of an encoder that learns a representation of human motion dynamics, an actor that proposes pedestrian displacements, a critic that evaluates imagined crowd trajectories, and a planner that selects actions. We first learn human motion dynamics through trajectory prediction on real-world pedestrian videos and then freeze the encoder to preserve them. Using this representation, the actor generates imagined crowd trajectories through repeated state updates, and the planner combines the critic's scores with user costs to select actions. Repeated planning advances the simulated crowd, while additional user costs introduce new control objectives without retraining. We extensively evaluate crowd generation under varied agent arrival conditions and run-time control across avoidance and attraction scenarios. Ctrl-CWM outperforms the state-of-the-art method on most crowd realism and collision metrics, and adapts crowd behaviors to user-specified objectives introduced during simulation. The project page is available at https://jungyu0413.github.io/Ctrl-CWM

関連論文

PR本紙発行元 EmplifAI