日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
arXiv:2507.19196

Towards Multimodal Social Conversations with Robots: Using Vision-Language Models

Towards Multimodal Social Conversations with Robots: Using Vision-Language Models

シェア:XThreadsFacebookLINEはてブBluesky

著者: Ruben Janssens, Tony Belpaeme

分類: cs.RO, cs.CL, cs.HC

原文アブストラクト

Large language models have given social robots the ability to autonomously engage in open-domain conversations. However, they are still missing a fundamental social skill: making use of the multiple modalities that carry social interactions. While previous work has focused on task-oriented interactions that require referencing the environment or specific phenomena in social interactions such as dialogue breakdowns, we outline the overall needs of a multimodal system for social conversations with robots. We then argue that vision-language models are able to process this wide range of visual information in a sufficiently general manner for autonomous social robots. We describe how to adapt them to this setting, which technical challenges remain, and briefly discuss evaluation practices.