日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
動作生成arXiv:2609.39575

ECHO-G: 音声に同期したヒューマノイド全身動作生成

ECHO-G: Embodied Co-speech Humanoid mOtion Generation

シェア:XThreadsFacebookLINEはてブBluesky

音声とタイミング付きテキストを条件に、拡散トランスフォーマでヒューマノイドロボットの全身発話動作を直接生成するフレームワークを提案し、実機展開まで示した。

著者: Yizhao Li, Pusen Gao, Ming Wang, Shaojie Shen, Shuo Yang, Hao Xu

分類: cs.RO, cs.AI

原文アブストラクト

Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts. Its Speech-Grounded Diffusion Transformer (SGDiT) combines frame-aligned acoustic features with token-level linguistic context, preserving their distinct granularities. Trained with rectified flow matching, it models one-to-many utterance-motion relationships directly in robot space. To support training and evaluation, we introduce a BEAT2-derived audio-text-robot dataset and a benchmark covering co-speech characteristics, robot-motion quality, and runtime efficiency. Comparative evaluation supports direct robot-space generation over the evaluated human-motion generation and retargeting pipelines, while modality ablations highlight the benefits of joint audio-text conditioning. We further demonstrate deployment on a physical humanoid robot. A complementary video-rating study also favors joint conditioning over the alternatives. The dataset and training, inference, and evaluation code are available through our project page.

関連論文

PR本紙発行元 EmplifAI