日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
知覚/ロボット盲導犬arXiv:2610.03187

ロボット盲導犬のための軽量・省リソースな知覚システム

Lightweight and Resource-Efficient Perception for Robotic Guide Dogs

シェア:XThreadsFacebookLINEはてブBluesky

360度カメラと2D LiDARを融合し、歩行者の視点で障害物を認識・説明するオンデバイス知覚モジュールを開発。55W以下でリアルタイム動作し、GuideDogQAベンチマークで83.8%の精度を達成した。

詳しい要約

1. どんなもの?

- ロボット盲導犬向けの軽量・省リソースな知覚システム - 360 cameraと2D LiDARを融合し、衝突回避と人間中心のガイダンスを実現 - 移動物体の検出・追跡、および歩行不可能時のvision-language modelによる経路説明を含む - オンデバイスで動作し、実世界egocentric GuideDogQAベンチマークで83.8%の精度を達成

2. 先行研究と比べてどこがすごい?

- 従来はカメラや2D/3D LiDARの生センサデータに依存し、距離測定に重点 - それらはロボット中心の衝突回避や安全性には有効だが、人間中心のガイダンスには不向き - 本研究は障害物の種類と関連性を認識し、ユーザーとの空間関係を明確かつ行動可能な形で説明 - 360 cameraと2D LiDARを融合し、移動物体検出・追跡を統合 - 歩行不可能時にvision-language modelで経路説明を提供し、ユーザーの不安を軽減

3. 技術・手法の肝は?

- 360 cameraと2D LiDARの融合による深度推定と近距離知覚 - 移動物体の検出・追跡を組み合わせた人間中心ガイダンス - 歩行不可能な状況ではvision-language modelが経路説明を生成 - 完全オンデバイスで動作し、リアルタイム性能を55 W以下で維持

4. どうやって有効だと検証した?

- 融合した360 camera-LiDAR深度の検証で、信頼できる近距離知覚を確認 - ただし中距離に固有のバイアスがあることも判明 - システム全体が55 W以下でリアルタイム性能を維持 - 実世界egocentric GuideDogQAベンチマークで83.8%の精度を達成(GPT-4oは67.1%)

5. 議論はある?

- 融合深度は近距離では信頼できるが、中距離に固有のバイアスが存在 - システムは55 W以下でリアルタイム動作を維持し、実用性を示す - 四足歩行ロボット上でも人間中心のガイダンスが実現可能であることを実証 - 具体的な限界や課題については要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:GPT-4o - 関連手法:360 cameraと2D LiDARの融合、vision-language model、移動物体検出・追跡 - 同分野の定番:LiDAR-based perception、vision-language navigation、egocentric vision for assistive robotics

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jinse Kwon, Yoojin Lim, Choonghan Lee, Yongseung Yu, Yongin Kwon, Jemin Lee

分類: cs.DC, cs.CV

原文アブストラクト

Robotic guide dogs should understand their surroundings, objects, and potential risks. Prior research has focused on raw sensor data from cameras and 2D or 3D LiDAR, which precisely measure distance points rather than provide a semantic understanding of the scene. While these physical measurements are effective for robot-centric collision avoidance and robot safety, they are not suitable for human-centric guidance. The system should recognize the type and relevance of obstacles and explain them, clearly and actionably, in terms of their spatial relation to the user. We present complete on-device perception modules that fuse a 360 camera and a 2D LiDAR for reliable collision avoidance, with moving-object detection and tracking for human-centric guidance. Finally, in walking-impossible situations, a vision--language model delivers pathway explanations as a safety mechanism to reduce user anxiety. In experiments, verification of fused 360 camera--LiDAR depth shows reliable near-range perception but inherent mid-range bias, while the system as a whole sustained real-time performance under 55 W. On the real-world egocentric GuideDogQA benchmark, our system achieved 83.8\% accuracy, compared with 67.1\% for GPT-4o. These results demonstrate that practical human-centric guidance with real-time on-device inference is feasible even on quadrupeds.

PR本紙発行元 EmplifAI