社会的直感と機械推論:複数モダリティからの人とロボットの相互作用予測
Social Intuition vs. Machine Reasoning: Anticipating Human-Robot Interaction from multiple modalities
サービスロボットの視点から人が相互作用する意図を予測するタスクで、人間の直感と機械学習モデル(ポーズベースモデルや視覚言語モデル)の性能を比較し、人間が特にフルビデオ入力で優れることを示した。
著者: Raphael Lorenzo-Louis, Bertrand Luvison, Serena Ivaldi
分類: cs.CV
原文アブストラクト
Anticipating whether a person will interact from one's own perspective is a highly intuitive task for humans, that relies on a combination of cues. We investigate how humans perform at predicting a person's intention to interact from a service robot's point of view, using pose-only or full video input, then benchmark different lightweight pose-based models and state-of-the-art vision-language models. We conducted our benchmark on the HUI360 dataset on a fixed pilot subset of 100 test tracks (25 positive, 75 negative). We found that with pose-only input, human annotators outperform lightweight trained pose models but not by large margins (+0.08 in F1-Score). But when given full egocentric video with a target bounding box, human annotators perform substantially better and largely outperform the Vision-Language Models (+0.2 in F1-Score). We also compared VLMs of different size and under different input conditions, and found that the best results do not correlate with model size. Our result confirms that predicting interactions is a challenging task for social robots and that reasoning-capable models are necessary but their actual reasoning capabilities alone do not suffice to match the social intuition of humans.