日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
音響ナビゲーションarXiv:2609.21792

AcousticDiffusion: 意味条件付き音響誘導拡散ポリシーによる救助支援

AcousticDiffusion: Semantically Conditioned Audio-Guided Diffusion Policy for Search-and-Rescue Assistance

シェア:XThreadsFacebookLINEはてブBluesky

音声認識とマイクアレイの到来方向推定を統合し、拡散モデルで経路を生成する音響誘導ナビゲーション手法を提案。四足歩行ロボットに実装し、救助者への接近性能を向上させた。

詳しい要約

1. どんなもの?

- 視界が劣化・遮蔽された環境で救助ロボットが人間の呼び声へ向かうための、意味論的条件付き音響誘導 diffusion policy『AcousticDiffusion』。 - 凍結済み音声認識器で 10.24 s 窓を処理し、speech gating と distress-aware prioritization で音源レベルの navigation role に変換。 - microphone-array の direction-of-arrival を再帰的にロボット中心の Bayesian bird's-eye-view belief field へ統合。 - ego-motion compensation で連続観測を整合し、方位由来の range uncertainty を保ちつつ音源位置を制約。 - 意味 belief・直近の音響観測・audio features・robot state を条件に diffusion model が waypoint 軌道を生成する。

2. 先行研究と比べてどこがすごい?

- 古典的 planner の A* や RRT と比較し、ZSL-1 quadruped 上で再学習なしに mean bearing error 64.9° を達成(A* 98.2°、RRT 90.4°)。 - ODAS(Open embedded Audition System)由来の guidance を使う古典 planner の mean final source distance 3.96 m に対し 2.48 m へ 37.4% 改善。 - 不完全な音響 localization 下でも、不確実な音響観測をより近い接近へ変換できる点を示した。 - 意味論的条件付けにより HELP 指定話者を競合話者より 91.07% の窓で優先するなど、単なる音源定位を超えた役割選択を実現。

3. 技術・手法の肝は?

- 凍結済み pretrained audio recognizer が 10.24 s 窓を処理。 - speech gating と distress-aware prioritization で認識出力を source-level navigation role へ変換。 - microphone-array の direction-of-arrival 測定を再帰的に robot-centric Bayesian bird's-eye-view belief field へ統合。 - ego-motion compensation で連続観測を整合し、方位由来の range uncertainty を保持しつつ音源位置を漸進的に制約。 - semantic belief・recent acoustic observations・audio features・robot state を条件として diffusion model が waypoint trajectories を生成。

4. どうやって有効だと検証した?

- 録音音声を用いた synthetic-navigation validation set で評価。 - mean end-point bearing error 11.20°、軌道の 91.78% が呼び手の 30° 以内に整列。 - distractor rejection は 89.20%〜98.99%、HELP 指定話者を競合話者より優先した割合は 91.07%。 - ZSL-1 quadruped 上で追加再学習なしにオンライン展開し、mean bearing error 64.9°(A* 98.2°、RRT 90.4°)、mean planner compute time 6.07 ms。 - mean final source distance は ODAS 由来 guidance の古典 planner 3.96 m から 2.48 m へ 37.4% 改善。

5. 議論はある?

- 音響 localization は不完全であるにもかかわらず、不確実な音響観測をより近い接近へ変換できることを示す。 - 課題として、実機での mean bearing error 64.9° は synthetic 検証の 11.20° より大きく、sim-to-real のギャップが示唆される。 - 要旨からは、計算コスト・失敗事例・倫理面・限界の詳細な議論は不明。

6. 次に読むべき論文は?

- A* や RRT などの classical planners。 - ODAS(Open embedded Audition System)。 - diffusion policy 系の手法。 - audio-guided navigation や search-and-rescue robotics の関連研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Iana Zhura, Didar Seyidov, Dmitrii Plotnikov, Hajira Amjad, Miguel Altamirano Cabrera, Dzmitry Tsetserukou

分類: cs.RO

原文アブストラクト

Navigating toward human callers is an important capability for rescue robots operating where visual contact is degraded or occluded. We present AcousticDiffusion, a semantically conditioned, audio-guided diffusion policy for human-directed navigation. A frozen pretrained audio recognizer processes 10.24 s windows, with speech gating and distress-aware prioritization converting recognition outputs into source-level navigation roles. Microphone-array direction-of-arrival measurements are recursively integrated into a robot-centric Bayesian bird's-eye-view belief field. Ego-motion compensation aligns successive observations, progressively constraining source position while preserving bearing-induced range uncertainty. The semantic belief, recent acoustic observations, audio features, and robot state condition a diffusion model that generates waypoint trajectories. On a synthetic-navigation validation set using recorded audio, AcousticDiffusion achieves a mean end-point bearing error of 11.20 degrees, with 91.78% of trajectories aligned within 30 degrees of the caller. Distractor rejection ranges from 89.20% to 98.99%, and the policy favors a HELP-designated caller over a competing speaker in 91.07% of windows. Deployed online on a ZSL-1 quadruped without additional retraining, it achieves a mean bearing error of 64.9 degrees, compared with 98.2 degrees for A* and 90.4 degrees for RRT, with a mean planner compute time of 6.07 ms. Despite imperfect acoustic localization, the reported mean final source distance is reduced from 3.96 m for the classical planners using ODAS-derived (Open embedded Audition System) guidance to 2.48 m, a 37.4% improvement. These results demonstrate the framework's ability to translate uncertain acoustic observations into closer approaches to human callers.

PR本紙発行元 EmplifAI