日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
身体性会話AI/表情生成arXiv:2609.33095

REALM: 身体性を持つ反応的聴き手のための粗から細への生成フレームワーク

REALM: A Coarse-to-Fine Generative Framework for Embodied Reactive Listening

シェア:XThreadsFacebookLINEはてブBluesky

音声に反応する聴き手の顔動作を生成するため、話者音声と聴き手の動作履歴を遅延を考慮して融合し、粗い動作軌道を確率的な表情残差で精緻化するフレームワークを提案した。

著者: Peizhen Li, Longbing Cao, Yang Zhang

分類: cs.RO

原文アブストラクト

Generating responsive listener facial motion is an important task for embodied conversational AI. Two modeling challenges are central: accounting for the timing of speaker cues while maintaining continuity with the listener's ongoing motion, and capturing locally variable facial events alongside the overall motion trajectory. Listener responses may follow preceding cues with a temporal lag, while brief expressions and blinks introduce variation that is difficult to predict deterministically. These challenges motivate a framework that combines history-aware temporal alignment with stochastic expression refinement. We propose REALM (Reactive Embodied Audio-driven Listening Model), a coarse-to-fine framework for audio-driven reactive listening. A Reactive Gated Speaker-Listener Fusion module combines listener motion history with speaker audio through a delay-centered attention prior and adaptive gating. A coarse decoder predicts a base motion trajectory, which is augmented by audio-conditioned stochastic residuals in the expression subspace while retaining the coarse pose parameters. Evaluations on ViCo and L2L show improvements over the evaluated baselines across multiple motion-quality metrics. Additional analyses examine delay sensitivity, gate behavior, and blink dynamics. Finally, deployment on an Ameca humanoid robot and a perceptual user study demonstrate the applicability of the generated behavior to physical embodiment. Code: https://github.com/lipzh5/REALM Demo: https://youtu.be/Tf5mpd5S8VQ

PR本紙発行元 EmplifAI