日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ナビゲーションarXiv:2609.40177

Social-WM: ロボットのソーシャルナビゲーションのための安全認識型潜在ワールドモデル

Social-WM: Safety-Aware Latent World Models for Robot Social Navigation

シェア:XThreadsFacebookLINEはてブBluesky

自己中心的なRGB動画から潜在ワールドモデルを学習し、指令された行動と実際に実行可能な行動の乖離を安全性シグナルとして活用することで、歩行者との衝突を抑えたソーシャルナビゲーションを実現した研究。

著者: Zhihao Zheng, Mooi Choo Chuah

分類: cs.RO

原文アブストラクト

Safe social navigation requires a robot to anticipate not only the future consequences of its actions, but also whether a nominal action can actually be executed under surrounding physical and social constraints. We present Social-WM, an efficient latent world-model planning framework trained from egocentric RGB video sequences. Our key observation is that social-navigation experience contains a systematic discrepancy between the nominal action and the realizable action: a nominal forward action may be fully executed in free space, but needs to be constrained when heading towards a pedestrian or obstacle. Social-WM learns these safety-relevant consequences directly through action-conditioned future prediction, where the target is the actual observed future following each command. We further introduce a realizable inverse-dynamics objective that associates observed latent transitions with the action actually realized rather than the nominal one. At deployment, candidate actions are imagined through the latent world model, and the inverse dynamics model estimates their realizability; nominal--realizable discrepancy then provides a safety signal before execution. The learned dynamics and realizability model remain goal-independent and support both position- and image-goal navigation. On Social-HM3D, Social-WM achieves 63.77% success while reducing human collisions to 21.67%, and maintains strong performance under zero-shot transfer to Social-MP3D, without explicit pedestrian tracking, privileged human state, or online reinforcement learning.

関連論文

PR本紙発行元 EmplifAI