日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
動作生成arXiv:2609.37708

生成的相互作用:二層潜在ダイナミクスによる多人数人体動作の織り成し

Generative Interactions: Weaving Multiparty Human Motion with Bilevel Latent Dynamics

シェア:XThreadsFacebookLINEはてブBluesky

集団レベルの潜在状態と個人レベルの潜在状態を階層的に組み合わせ、多人数の社会的動作生成をメタ転移学習として定式化したモデルBRAIDを提案。

詳しい要約

1. どんなもの?

複数人の社会的行動を生成するための階層的逐次潜在変数モデル BRAID を提案する論文。 - 社会的運動生成を meta-transfer learning 問題として定式化。 - シーンを group-level 潜在状態と person-level 潜在状態で表現。 - 完全・疎・部分観測下での coherent な生成を可能にする。 - 下流の embodied-agent システムへのインターフェースとなる social-state ベクトルを提供。

2. 先行研究と比べてどこがすごい?

既存の社会的運動モデルはもっともらしい軌道を優先し、interaction state を暗黙のままにしていた。 - そのためグループ・タスク・部分観測への転移が制限されていた。 - BRAID は interaction state を明示的に階層潜在変数としてモデル化。 - 共有 interaction prior をデータセット間で学習し、任意の context set で適応可能。 - これにより転移性と部分観測への頑健性を向上。

3. 技術・手法の肝は?

Bilevel Representations for Agent Interaction Dynamics (BRAID) を導入。 - 階層的逐次潜在変数モデルで、group-level と person-level の潜在状態を持つ。 - group-level 潜在状態が共有 interaction dynamics を捕捉。 - person-level 潜在状態が進化する group context に条件付けられた個々の行動を捕捉。 - meta-transfer learning として、共有 interaction prior を学習し、観測された人と関節の任意の context set で適応。 - SMPL-based 表現で統一。

4. どうやって有効だと検証した?

SMPL-based 表現の下で評価。 - social forecasting、tracking and in-filling、response generation のタスクで検証。 - 再構成精度だけでなく、realism、diversity、temporal alignment、interpersonal coordination を評価する指標を使用。 - 階層潜在空間を分析し、group-level と individual-level の分離可能な構造を捕捉することを示す。

5. 議論はある?

階層潜在空間が group-level と individual-level の構造を分離可能に捕捉することを分析。 - 完全・疎・部分観測下での coherent な生成が可能。 - コンパクトな social-state ベクトルが下流の embodied-agent システムのインターフェースになり得る。 - 限界や課題については要旨からは不明。

6. 次に読むべき論文は?

要旨で参照・比較されている研究は明示されていない。 - 関連手法として social motion models、meta-transfer learning、hierarchical latent-variable models、SMPL-based human motion representation が挙げられる。 - 同分野の定番として social forecasting、tracking、response generation に関する研究を読むと良い。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ojas Shirekar, Yash Surange, Agustinas Jučas, Chirag Raman

分類: cs.AI, cs.LG

原文アブストラクト

Human social behaviour is not a collection of independent motions, but a jointly organised process in which group dynamics and individual variation continuously shape one another. Yet existing social motion models often prioritise plausible trajectories while leaving interaction state implicit, limiting their ability to transfer across groups, tasks, and partial-observation regimes. To address this gap, we introduce Bilevel Representations for Agent Interaction Dynamics (BRAID), a hierarchical sequential latent-variable model for generative multi-person interaction. BRAID explicitly formulates social motion generation as a meta-transfer learning problem: shared interaction priors are learned across datasets and adapted through arbitrary context sets of observed people and joints. The model represents each scene through a group-level latent state that captures shared interaction dynamics and person-level latent states that capture individual behaviour conditioned on the evolving group context. This modelling choice enables coherent generation under full, sparse, or partial observations while exposing compact social-state vectors that can serve as an interface for downstream embodied-agent systems. We evaluate BRAID under a unified SMPL-based representation on social forecasting, tracking and in-filling, and response generation, using metrics that assess not only reconstruction accuracy but also realism, diversity, temporal alignment, and interpersonal coordination. We further analyse the hierarchical latent space, showing that it captures separable group- and individual-level structure.

関連論文

PR本紙発行元 EmplifAI