日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マルチモーダル参照理解arXiv:2510.08278

奥行きを考慮したマルチモーダル身体参照理解手法

A Multimodal Depth-Aware Method For Embodied Reference Understanding

シェア:XThreadsFacebookLINEはてブBluesky

言語指示と指差しの手がかりから対象物体を特定する身体参照理解において、LLMによるデータ拡張と深度マップ、深度認識決定モジュールを組み合わせ、曖昧な場面での精度を向上させた。

著者: Fevziye Irem Eyiokur, Dogucan Yaman, Hazım Kemal Ekenel, Alexander Waibel

分類: cs.CV, cs.HC, cs.RO

原文アブストラクト

Embodied Reference Understanding requires identifying a target object in a visual scene based on both language instructions and pointing cues. While prior works have shown progress in open-vocabulary object detection, they often fail in ambiguous scenarios where multiple candidate objects exist in the scene. To address these challenges, we propose a novel ERU framework that jointly leverages LLM-based data augmentation, depth-map modality, and a depth-aware decision module. This design enables robust integration of linguistic and embodied cues, improving disambiguation in complex or cluttered environments. Experimental results on two datasets demonstrate that our approach significantly outperforms existing baselines, achieving more accurate and reliable referent detection.