社会的ロボット仲介のためのマルチモーダル音声活動予測:期待される行動と展開制約
Multimodal Voice Activity Projection for Social Robot Mediation: Expected Behavior and Deployment Constraints
音声と視覚の同期情報から会話のターンテイキングを予測し、ロボットが仲介者として取るべき行動(待機・傾聴・介入など)を導出する知覚レイヤーを提案。実時間推論やマルチモーダル同期などの展開制約も議論する。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Antonio Cano, Guillermo Perez, Luis Merino, Randy Gomez
分類: cs.RO
原文アブストラクト
Turn-taking prediction is especially relevant for social robots that act as mediators in human-human interaction, where the expected action is often not to speak, but to orient, wait, avoid interruption, or prepare a balanced intervention. This paper presents Multimodal Voice Activity Projection (MM-VAP) as a human state-aware perception layer for future robot mediation behavior. The model estimates the future evolution of the conversational floor from synchronized audio-visual evidence and derives turn-taking events such as Hold, Shift, Shift prediction, Backchannel prediction, and overlap-related states. The approach uses VA-related pretrained audio-visual encoders, LoRA adaptation, inter-speaker attention, and zero-shot event inference from future voice activity projections. Experiments on NoXi, NoXi+J, and Haru EDR support the feasibility of this formulation, especially for floor management events that can be connected to gaze preparation, active listening, and conservative intervention. Finally, the paper defines the expected robot output interface and discusses the main deployment constraints, including real-time inference, preprocessing latency, multimodal synchronization, and input-quality monitoring.