日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ソーシャルロボットarXiv:2608.16686v1

感情ループを閉じる:マルチモーダルな話者・聞き手の感情ダイナミクスを考慮した共感型ソーシャルロボット

Closing the Affective Loop: Multimodal Speaker-Listener Emotion-Dynamics-Aware Empathetic Social Robots

シェア:XThreadsFacebookLINEはてブBluesky

ユーザーの発話内容だけでなく、対話中の感情の動的な変化にも応答する共感型ロボットを提案。Misty IIロボットに実装し、話者と聞き手の感情状態を追跡して応答生成に活用するシステムを評価した。

詳しい要約

1. どんなもの?

AffectLoopは、Misty IIロボット上に実装されたマルチモーダルな対話システムで、話者の言語・表情の感情ダイナミクスと、ロボット聞き手自身の言語・行動の感情状態を推定し、LLMベースの応答生成に両方の感情ストリームを条件付けます。ロボットは短い共感応答と感情的に一致した身体動作を生成し、話者-聞き手の感情ループを閉じます。

2. 先行研究と比べてどこがすごい?

既存の共感対話システムはテキスト中心で、ユーザーの感情からシステム応答への一方向マッピングとして共感をモデル化することが多く、身体化された話者-聞き手の感情交換を捉えることができません。AffectLoopは、話者の感情ダイナミクスと聞き手の感情状態の両方を明示的にモデル化し、マルチモーダルな感情ループを形成する点で優れています。

3. 技術・手法の肝は?

手法の肝は、話者の言語・表情の感情ダイナミクスを追跡し、ロボット聞き手自身の言語・行動の感情状態を推定し、LLMベースの応答生成を両方の感情ストリームに条件付けることです。これにより、応答は話者の感情変化に適応し、聞き手の感情状態も反映されます。

4. どうやって有効だと検証した?

5人の参加者によるパイロット被験者内研究で、話者・聞き手の感情状態入力を省略した同一のベースラインと比較しました。提案システムは全体的な印象評価が高く、特に共感応答とユーザー満足度で優れていました。また、ログ分析により、話者-聞き手の感情的一致と、感情価に基づく苦痛回復の強化が示されました。

5. 議論はある?

議論としては、サンプルサイズが小さく(5人)、パイロット研究であるため、結果の一般化には限界があります。また、感情ダイナミクスの推定精度や、LLM生成の応答品質が結果に与える影響についての詳細な分析は要旨からは不明です。

6. 次に読むべき論文は?

要旨で参照されている研究は明示されていませんが、関連する分野として、共感対話システム、マルチモーダル感情認識、社会的ロボットの研究が挙げられます。具体的には、LLMベースの対話システムや、感情状態を考慮した応答生成に関する論文が次に読むべき候補です。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zi Haur Pang, Casey Kennington, Tatsuya Kawahara

分類: cs.HC, cs.CL, cs.RO

原文アブストラクト

Empathetic social robots should respond not only to what users say, but also to how their emotions dynamically evolve during interaction. However, existing empathetic dialogue systems are often text-centered and primarily model empathy as a one-way mapping from the user's emotion to the system response, limiting their ability to capture embodied speaker--listener affective exchange. We present AffectLoop, a multimodal speaker-listener emotion-dynamics-aware spoken dialogue system implemented on the Misty II robot. The system tracks the speaker's verbal and facial affective dynamics, estimates the robot listener's own verbal and behavioral affective state, and conditions LLM-based response generation on both affective streams. The robot then generates a short spoken empathetic response together with emotionally congruent embodied behavior, forming a closed speaker--listener affective loop. We evaluate the system in a pilot within-subject study with five participants, comparing it with an otherwise identical utterance-conditioned baseline that omits the speaker- and listener-affective-state inputs. The proposed system received higher overall impression ratings, especially for empathetic response and user satisfaction. Post-hoc log analysis further showed higher speaker-listener affective alignment and stronger valence-based distress recovery. These preliminary results suggest that explicitly modeling both speaker emotional dynamics and listener affective state can improve embodied empathetic interaction.