MIRA: 身体性を備えたコンパニオン向けリアルタイム全二重ヒューマンロボットインタラクション
MIRA: Real-Time Full-Duplex Human-Robot Interaction for Embodied Companions
ストリーミング音声から応答と身体動作を同時生成し、割り込み可能な動作をリアルタイムで実行する全二重対話フレームワークを提案し、ヒューマノイドロボットに実装した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Lijian Lin, Ye Zhu, Fan Zhang, Yunfei Liu, Baofeng Li, Xianwen Zeng, Jianan Wang, Yu Li
分類: cs.RO
原文アブストラクト
Real-time embodied companion interaction requires a robot to infer user intent from streaming speech, generate timely responses, and execute expressive, interruptible motions. Existing systems typically decouple dialogue orchestration from gesture synthesis, relying on offline motion generation from complete audio. This separation leaves open how a deployed robot can dynamically synchronize response content, prosodic timing, and physical safety under incremental inputs and uncertain turn boundaries. We present MIRA, a unified framework for full-duplex embodied companion interaction. Given streaming user speech, dialogue history, and vocal affect, MIRA predicts both the response text and an explicit embodiment cue. Discrete social behaviors (\eg listening, greeting) are mapped to validated robot trajectories, while open-ended speaking is paired with streaming, co-speech motion. This generative motion is governed by a predict-more-than-commit sliding window that provides temporal look-ahead for motion continuity while limiting physical commitment to a short, cancellable prefix. Crucially, we design CORTEX, a dual-timescale interaction policy that manages low-latency streaming and deliberative turn decisions, backed by a robot-side execution layer that enforces physical safety constraints at the control rate. We deploy MIRA on an Astribot S1 humanoid robot. Quantitative evaluations demonstrate competitive audio-motion alignment relative to state-of-the-art motion-generation baselines, while real-robot deployment measurements characterize streaming responsiveness and interruption handling.
関連論文
- 個々の特性を考慮した支援型ヒューマンロボットインタラクションにおけるエンゲージメントと侵入性の理解ヒューマンロボットインタラクション
- ロボットの声の高さは子どものストレスを和らげるか?ヒューマンロボットインタラクション
- OmniAI: 人間とドローンの対話のための表面適応型空中投影インターフェースヒューマンロボットインタラクション
- 一目でわかる:顔から見た目の性格を推定するヒューマンロボットインタラクション
- ロボットのためのインテリジェントクラウドエッジマルチモーダル対話システムヒューマンロボットインタラクション
- ロボットによる歩行案内中の高齢者への触覚的接触の影響ヒューマンロボットインタラクション