ロボット用基盤モデルはどう選ぶべきか?ソーシャルロボットのためのコミュニティ評価フレームワークを支持する
How Should I Pick a Foundation Model for My Robot? In Favor of a Community Evaluation Framework for Social Robots
ソーシャルロボット向け基盤モデルの選択を支援するため、5つの評価次元と3段階の評価パラダイムを提案し、コミュニティでの共同構築を呼びかける論文。
著者: Eric Nichols, Alva Markelius, Hatice Gunes
分類: cs.RO, cs.CL, cs.HC
原文アブストラクト
Researchers who seek to build social robot applications on foundation models are faced with a difficult question: how should we pick a model? Public leaderboards offer little guidance: the demands of real-time, embodied social interaction lie largely outside their focus. And direct evaluation is impractical at scale: each embodied study requires scarce participant, robot, and experimenter time. In this paper, we identify five evaluation dimensions for foundation models in social robots: (i) conversational competence, (ii) user safety, (iii) embodied character, (iv) target scene effectiveness, and (v) audience appropriateness. To make model selection cheaper and better informed, we propose a three-tiered evaluation funnel paradigm that first filters with general metrics, then extends to simulated interactions, and terminates in more expensive, robot-specific evaluation. We map all five dimensions across all three tiers, chart where applicable evaluation methods exist and are missing, and close with a call to action: let's build the evaluation framework together as a community.