社会的ロボットとの音声対話のための拡張可能なベンチマークに向けて
Towards an Extensible Benchmark for Spoken Dialogue with Social Robots
ロボットと人間の音声対話における課題(聞き返し、割り込み、身体信号、時間制約)を評価するベンチマークを提案し、その拡張ビジョンとリアルタイム通信フレームワークReticoを示す。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Casey Kennington, Ross Mead, Saad Elbeleidy, Jesse Thomason
分類: cs.RO
原文アブストラクト
Language models provide a plug-and-play interface between humans and robots, but important challenges remain when speech, dialogue, fast interaction, and collaboration are required. We propose a benchmark for the community to use as a way to explore common spoken dialogue artifacts between robots and humans, including requests for clarification, interruptions, embodied signals (e.g., head nods or facial cues), and time constraints. We also explain our vision to extend the benchmark for other aspects of human-robot interaction that are important to the larger research community. To facilitate the benchmark, we further propose using \textit{Retico}, a real-time communication framework that fulfills important technical requirements to enable robots to have spoken dialogue capabilities.