日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
群制御arXiv:2609.27816

VLM-LLM推論と到達可能性解析による安全なマルチロボット協調

Safe Multi-Robot Coordination via VLM-LLM Reasoning and Reachability Analysis

シェア:XThreadsFacebookLINEはてブBluesky

視覚付き四足歩行ロボットとカメラなし車両の異種チームで、VLMが環境を意味解釈しLLMがタスクを割り当て、ゾノトープ到達可能性ゲートで安全性を検証して協調ナビゲーションを実現する。

詳しい要約

1. どんなもの?

- 異種ロボットチーム(視覚付き四足歩行ロボットとカメラなし車両)の協調ナビゲーションのための集中型安全フレームワーク。 - 目標領域への移動中に静的・動的障害物を回避し、ロボット間の不安全な相互作用を防ぐ。 - 共有知覚の原則に基づき、視覚ロボットがMQTTブローカー経由でセマンティック環境情報を提供。 - VLMが視覚ストリームを解釈し、LLMが高レベルタスク割り当てを提案。 - ロボット固有のzonotope到達可能性ゲートが物理コマンド権限を制限。

2. 先行研究と比べてどこがすごい?

- 異種ロボットシステムにおける安全な協調を、VLM-LLM推論と到達可能性解析を組み合わせて実現。 - 従来の手法と比較して、セマンティック推論と形式的検証されたモーター実行を厳密に分離。 - カメラなしロボットでも共有知覚によりナビゲーション可能。 - 具体的な先行研究との比較は要旨からは不明。

3. 技術・手法の肝は?

- VLMが視覚ストリームを解釈し、セマンティックデータを保守的なメートル幾何制約(膨張障害物セット、安全コリドー、目標領域)にマッピング。 - LLMが高レベルタスク割り当てを提案するが、物理コマンド権限はロボット固有のzonotope到達可能性ゲートに制限。 - 検証エンジンが独立した到達可能チューブを伝播し、障害物回避、安全コリドー包含、ロボット間分離述語を評価してからコマンドを承認。 - 集中型サーバーがMQTTブローカー経由で共有知覚を提供。

4. どうやって有効だと検証した?

- クリアパスと動的障害物シナリオでのオンライン実験を実施。 - パイプラインが安全な動作を確実に承認し、制約違反時に保守的な再計画または保持動作をトリガーすることを確認。 - アドバイザリーなセマンティック推論と形式的に検証されたモーター実行の厳密なアーキテクチャ分離を実証。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、VLM、LLM、zonotope到達可能性解析、MQTT、共有知覚、異種マルチロボット協調が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Mohamed Dwedar, Ahmad Hafez, Alexander Jesser, Amr Alanwar

分類: cs.RO, cs.AI, eess.SY

原文アブストラクト

Safe coordination in heterogeneous machine-to-machine (M2M) robotic systems is challenging when robots differ in sensing capabilities, environmental awareness, and motion execution roles. This paper presents a centralized safety-aware M2M framework for cooperative goal-directed navigation in a heterogeneous mobile robot team comprising a vision-capable quadruped and a camera-less robotic vehicle. The objective is to guide both platforms toward a goal region while avoiding static and dynamic obstacles and preventing unsafe inter-robot interactions. Under the principle of shared perception, the vision-capable robot provides semantic environmental awareness through a centralized server over an MQTT broker, enabling the camera-less platform to navigate using this shared scene representation alongside its own odometry, IMU, and state feedback. A vision-language model (VLM) interprets the visual stream, and the extracted semantic data is mapped into conservative metric geometric constraints, including inflated obstacle sets, safe corridors, and goal regions. A large language model (LLM) proposes high-level task allocations, while physical command authority is restricted to a robot-specific zonotope reachability gate. This verification engine propagates independent reachable tubes to evaluate obstacle avoidance, safe-corridor containment, and inter-robot separation predicates before approving commands. Online experiments across clear-path and dynamic-obstacle scenarios show that the pipeline reliably approves safe motion, triggers conservative replanning or holding maneuvers upon constraint violation, and enforces a strict architectural separation between advisory semantic reasoning and formally verified motor execution.

関連論文

PR本紙発行元 EmplifAI