ロボットのためのインテリジェントクラウドエッジマルチモーダル対話システム
An Intelligent-Cloud Edge Multimodal Interaction System for Robots
限られた計算資源のロボット向けに、改良したジェスチャ検出器とLLM/VLMエージェントを組み合わせたクラウドエッジ型マルチモーダル対話フレームワークを提案し、複雑環境での対話成功率を実証した。
分類: cs.RO, cs.AI
原文アブストラクト
Robust human-robot interaction in complex environments requires accurate gesture perception, semantic scene understanding, and reliable task planning under limited onboard computing resources. This paper presents a cloud-edge multimodal interaction framework that integrates an enhanced YOLO-based gesture detector with coordinated large language model (LLM) and vision-language model (VLM) agents. The proposed detector, incorporates the Convolutional Block Attention Module (CBAM) into the neck and replaces the baseline bounding-box regression objective with Distance-IoU (DIoU) loss. These modifications improve feature discrimination and localization for small or partially occluded gestures in complex backgrounds. The cloud layer performs gesture detection, scene understanding, multimodal fusion, and action planning, whereas the TonyPi robot locally handles data acquisition, communication, action execution, and feedback. Experiments on a public gesture dataset and a custom dataset show that YOLO-DC achieves precision values of 98.9% and 95.0%, with mAP@0.5 values of 90.7% and 92.7%, respectively. System-level evaluation yields success rates of 95%, 88%, and 82% for single-action, composite-action, and vision-dependent tasks. A 30 participant evaluation yields an overall mean satisfaction score of 3.69 out of 5. These results demonstrate the feasibility of combining refined gesture detection with multimodal agents for resource-constrained robotic interaction.
関連論文
- ロボットの声の高さは子どものストレスを和らげるか?ヒューマンロボットインタラクション
- OmniAI: 人間とドローンの対話のための表面適応型空中投影インターフェースヒューマンロボットインタラクション
- WCM: 汎用ヒューマンロボットインタラクションのための世界認知モデルヒューマンロボットインタラクション
- 博物館に戻る:感情シミュレーションの有無によるアンドロイド「アンドレア」の受容調査ヒューマンロボットインタラクション
- 一目でわかる:顔から見た目の性格を推定するヒューマンロボットインタラクション
- WCM: 汎用ヒューマンロボットインタラクションのための世界認知モデルヒューマンロボットインタラクション