インターネット動画から方策状態空間でソーシャルナビゲーションを学習
Learning Social Navigation from Internet Videos in the Policy State Space
単眼歩行動画を方策の状態空間における閉ループのソーシャルナビゲーション訓練環境に変換する手法を提案し、実機実験で高い成功率を達成した。
著者: Jiaming Wang, Duc Thang Nguyen, Jizhuo Chen, Volodymyr Shcherbyna, Diwen Liu, Zhengcheng Shen, Harold Soh
分類: cs.RO, cs.CV, cs.LG
原文アブストラクト
Training robust social-navigation policies requires simulators with diverse scene layouts, terrain, and human motion, but constructing such environments and specifying pedestrian behavior is costly. We propose an efficient pipeline that converts ordinary monocular walking videos directly into closed-loop social-navigation training environments in the policy's state space. Our key observation is that local social navigation primarily depends on two types of information: where the robot can traverse and how nearby pedestrians move. We therefore represent the static scene as a metric traversability map, which can be rigidly transformed under counterfactual robot motion, while directly replaying the pedestrian trajectories recovered from the video over time. This abstraction allows us to define the forward dynamics directly in the policy's state space and efficiently simulate counterfactual robot states without reconstructing or rendering photorealistic observations. The resulting policy achieves 81.2% success in the independent Arena benchmark, compared with 75.0% for the strongest baseline, and succeeds in 19/20 real-robot trials without policy fine-tuning.
関連論文
- どこに加わるべきか?言語誘導型目標予測によるロボットのグループ参加ソーシャルナビゲーション
- PISHYAR: 視覚障害者向けの社会的インテリジェントなスマート杖による屋内ソーシャルナビゲーションとマルチモーダルな人間-ロボットインタラクションソーシャルナビゲーション
- 障害物からエチケットへ:VLMに基づく経路選択によるロボットのソーシャルナビゲーションソーシャルナビゲーション
- 移動サービスロボットのためのペア単位の人間-人間インタラクション検出・認識フレームワークソーシャルナビゲーション
- 人間の動作予測品質が制約空間におけるソーシャルロボットナビゲーション性能をどう左右するかソーシャルナビゲーション
- LISN: VLMベース制御器調整による言語指示型ソーシャルナビゲーションソーシャルナビゲーション