日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
社会ナビゲーションarXiv:2609.24317

拡散ステアリングによる性能を維持した社会ナビゲーションのオンライン適応

Performance-Preserving Online Adaptation in Social Navigation via Diffusion Steering

シェア:XThreadsFacebookLINEはてブBluesky

拡散ポリシーを固定してノイズポリシーのみを強化学習で訓練するDSRLを提案し、複数シードの統合でベースポリシーを強化。性能を保ちつつ社会規範に適応する効率的な学習と実機検証を示した。

詳しい要約

1. どんなもの?

- 社会ナビゲーションにおけるオンライン適応手法の提案。 - 拡散政策を固定しノイズ政策のみを強化学習で訓練するDSRLを適用。 - 複数シードで訓練した拡散ベースRL政策を統合しベース政策を構築。 - 性能を保持しつつ展開環境で効率的に学習し、社会的慣習に適応。 - 物理ロボットでの有効性もハードウェアインザループシミュレーションで確認。

2. 先行研究と比べてどこがすごい?

- 従来の深層強化学習はシミュレーションのみで多様なシナリオやロボット動力学、社会的慣習を再現困難。 - 展開環境でのファインチューニングが有望だが、ベースモデルの性能保持が課題。 - 提案手法は拡散政策を固定しノイズ政策のみを訓練することで性能を保持。 - 複数シードの拡散ベースRL政策を統合し、学習性能を向上。 - 他の手法と比較して効率的な学習と性能保持を実現。

3. 技術・手法の肝は?

- 拡散政策を固定し、ノイズ政策のみを強化学習で訓練するDSRLを適用。 - 複数シードで訓練した拡散ベースRL政策を統合しベース政策を構築。 - これにより、ナビゲーションの主要目的(歩行者回避と目的地到達)を損なわずに適応。 - 社会的慣習への適応を通じた柔軟な行動制御を可能に。

4. どうやって有効だと検証した?

- 他の手法と比較し、提案手法が性能を保持しつつ効率的な学習を可能にすることを評価。 - 社会的慣習への適応による柔軟な行動制御を確認。 - ハードウェアインザループシミュレーションを通じて物理ロボットでの有効性を確認。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究や関連手法は明示されていない。 - 同分野の定番として、深層強化学習(deep reinforcement learning)や拡散政策(diffusion policy)に関する研究が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Haruto Nagahisa, Kohei Matsumoto, Yuki Hyodo, Ryo Kurazume

分類: cs.RO

原文アブストラクト

In social navigation, modeling the complex interactions between humans and robots is difficult, and deep reinforcement learning has therefore been actively studied. However, because simulation alone cannot fully reproduce diverse scenarios, robot dynamics, and the social conventions that vary across deployment environments, fine-tuning in the deployment environment is promising. In doing so, learning that preserves the base model's performance is required, so as not to compromise the primary objective of navigation, namely avoiding pedestrians and reaching the destination. In this study, we propose a method that applies diffusion steering via reinforcement learning (DSRL), which trains only the noise policy while keeping the diffusion policy fixed, thereby achieving learning that preserves performance. Furthermore, we integrate diffusion-based RL policies trained with multiple seeds to construct the base policy, improving learning performance. Our evaluation shows that, compared with other methods, the proposed method enables efficient learning while preserving performance, and we confirm flexible behavior control through adaptation to social conventions, as well as its effectiveness on a physical robot through hardware-in-the-loop simulation.

関連論文

PR本紙発行元 EmplifAI