日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ヒューマンロボットインタラクションarXiv:2609.24547

MIRA: 身体性を備えたコンパニオン向けリアルタイム全二重ヒューマンロボットインタラクション

MIRA: Real-Time Full-Duplex Human-Robot Interaction for Embodied Companions

シェア:XThreadsFacebookLINEはてブBluesky

ストリーミング音声から応答と身体動作を同時生成し、割り込み可能な動作をリアルタイムで実行する全二重対話フレームワークを提案し、ヒューマノイドロボットに実装した。

詳しい要約

1. どんなもの?

- 実時間のfull-duplexなembodied companion interactionを実現する統合フレームワークMIRAを提案。 - ストリーミング音声からユーザ意図を推論し、応答生成と表現力豊かで中断可能な動作を実行する。 - 対話オーケストレーションとジェスチャ合成を分離せず、応答内容・韻律タイミング・物理安全性を動的に同期する。 - Astribot S1 humanoid robotに実装。

2. 先行研究と比べてどこがすごい?

- 既存システムは対話オーケストレーションとジェスチャ合成を分離し、完全な音声からのオフライン動作生成に依存。 - MIRAはストリーミング入力と不確実なターン境界下で、応答内容・韻律タイミング・物理安全性を動的に同期可能。 - リアルタイムfull-duplex embodied companion interactionを統合フレームワークとして実現。

3. 技術・手法の肝は?

- ストリーミングユーザ音声、対話履歴、vocal affectから応答テキストと明示的embodiment cueを予測。 - 離散的社会行動(listening, greeting等)を検証済みロボット軌道にマッピング。 - オープンエンドな発話にはストリーミングco-speech motionをペアリング。 - predict-more-than-commit sliding windowで動作連続性のための時間的先読みを提供し、物理的コミットを短いキャンセル可能プレフィックスに限定。 - CORTEX: 低遅延ストリーミングと熟慮的ターン決定を管理するdual-timescale interaction policy。 - ロボット側実行層が制御レートで物理安全制約を強制。

4. どうやって有効だと検証した?

- Astribot S1 humanoid robotにMIRAを配備。 - 定量的評価でstate-of-the-art motion-generation baselinesと比較し、競争力のあるaudio-motion alignmentを示す。 - 実ロボット配備測定でストリーミング応答性と中断処理を特徴付ける。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- state-of-the-art motion-generation baselines(具体的名称は要旨からは不明) - 関連手法としてco-speech motion generation、full-duplex dialogue systems、embodied companion interactionの定番研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Lijian Lin, Ye Zhu, Fan Zhang, Yunfei Liu, Baofeng Li, Xianwen Zeng, Jianan Wang, Yu Li

分類: cs.RO

原文アブストラクト

Real-time embodied companion interaction requires a robot to infer user intent from streaming speech, generate timely responses, and execute expressive, interruptible motions. Existing systems typically decouple dialogue orchestration from gesture synthesis, relying on offline motion generation from complete audio. This separation leaves open how a deployed robot can dynamically synchronize response content, prosodic timing, and physical safety under incremental inputs and uncertain turn boundaries. We present MIRA, a unified framework for full-duplex embodied companion interaction. Given streaming user speech, dialogue history, and vocal affect, MIRA predicts both the response text and an explicit embodiment cue. Discrete social behaviors (\eg listening, greeting) are mapped to validated robot trajectories, while open-ended speaking is paired with streaming, co-speech motion. This generative motion is governed by a predict-more-than-commit sliding window that provides temporal look-ahead for motion continuity while limiting physical commitment to a short, cancellable prefix. Crucially, we design CORTEX, a dual-timescale interaction policy that manages low-latency streaming and deliberative turn decisions, backed by a robot-side execution layer that enforces physical safety constraints at the control rate. We deploy MIRA on an Astribot S1 humanoid robot. Quantitative evaluations demonstrate competitive audio-motion alignment relative to state-of-the-art motion-generation baselines, while real-robot deployment measurements characterize streaming responsiveness and interruption handling.

関連論文

PR本紙発行元 EmplifAI