日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
共話ジェスチャー/個人化arXiv:2609.13876

あなたと一緒に話す:再学習不要のロボット共話ジェスチャーの個人化

Co-Speech with You: Training-Free Personalization of Robot Co-Speech Gestures

シェア:XThreadsFacebookLINEはてブBluesky

約10秒の登録動作から話者スタイルを抽出し、再学習なしでロボットの共話ジェスチャーを個人化する手法を提案。NAOロボット上でも有効性を確認。

詳しい要約

1. どんなもの?

- 訓練不要で個人適応する co-speech gesture 生成の pipeline - 新規ユーザーの約10秒の enrollment motion から style embedding を抽出 - 抽出した style を条件に、その話者の gesture style を生成 - 構成要素 - frozen audio-conditioned diffusion prior - gesture style encoder - lightweight conditioning adapters - 付随して公開 - Quest 3 capture application - 10名の spontaneous co-speech motion dataset - 評価指標 - Style Recognition Accuracy (SRA) - Fréchet Gesture Distance (FGD) - 実機展開 - 生成 gesture を physical NAO robot に retarget - 5つの TTS voice…

2. 先行研究と比べてどこがすごい?

- 従来の co-speech gesture 個人化は model retraining を要することが多い - 本手法は training-free で新規ユーザーに適応 - frozen prior のままでは SRA 27.6% だが、本手法は 69.5% に改善 - 同時に motion quality を維持 (FGD 34.2 vs 34.8) - enrollment embedding を他人のものに置換すると SRA 11.4% に低下 - style embedding が話者固有情報を実際に担っていることを示す - 実機 NAO robot 展開でも改善が転移 - 5 TTS voices 設定で SRA 67.3% を保持 - 要旨からは不明 - 既存の training-based personalization との直接比較数値

3. 技術・手法の肝は?

- frozen audio-conditioned diffusion prior を土台に利用 - gesture style encoder を speaker identity 識別で事前学習 - その後、lightweight conditioning adapters と jointly refined - diffusion objective で最適化 - enrollment motion 約10秒から single forward pass で reusable style embedding を抽出 - 抽出した style embedding を conditioning として diffusion prior に注入 - 生成 gesture を physical NAO robot へ retarget - 要旨からは不明 - encoder/adapter の具体的な architecture や diffusion の詳細

4. どうやって有効だと検証した?

- held-out speakers で評価 - 指標 - Style Recognition Accuracy (SRA): personalization の指標 - Fréchet Gesture Distance (FGD): motion quality の指標 - 結果 - frozen prior: SRA 27.6% - 本手法: SRA 69.5%、FGD 34.2 (frozen prior は 34.8) - enrollment embedding を他人に置換: SRA 11.4% - 実機検証 - 生成 gesture を physical NAO robot に retarget - 5つの TTS voices で音声合成した設定でも SRA 67.3% を保持 - 要旨からは不明 - 被験者による主観評価や統計的有意差の有無

5. 議論はある?

- 訓練不要で新規ユーザーに適応できる点を主張 - style embedding の寄与を ablation 的に確認 - 他人の embedding に置換すると SRA が 11.4% まで低下 - motion quality を維持しつつ personalization を改善 - FGD 34.2 vs 34.8 - 実機 NAO robot と TTS voices 設定への転移を確認 - 要旨からは不明 - 限界、失敗事例、計算コスト、倫理的議論

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究 - frozen audio-conditioned diffusion prior - gesture style encoder - lightweight conditioning adapters - Style Recognition Accuracy (SRA) - Fréchet Gesture Distance (FGD) - Quest 3 capture application - NAO robot - TTS voices - 関連手法 - co-speech gesture generation - diffusion-based motion generation - speaker/style personalization - motion retargeting - 要旨に明示的な論文名は無いため、上記の手法・指標をキーワードに辿るのが妥当

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Bosong Ding, Selma Ancel, Giacomo Spigler, Murat Kirtay

分類: cs.RO

原文アブストラクト

Personal robots should adapt their co-speech gesture style to a new user without requiring model retraining. We present a training-free personalization pipeline that combines a frozen audio-conditioned diffusion prior with a gesture style encoder and lightweight conditioning adapters. The encoder is first trained to discriminate speaker identities and then jointly refined with the adapters using the diffusion objective, enabling a reusable style embedding to be extracted from approximately 10 seconds of enrollment motion through a single forward pass. To support this setting, we also release a Quest~3 capture application and a dataset of spontaneous co-speech motion from ten participants. We evaluate the system on held-out speakers using Style Recognition Accuracy (SRA) and Fr'echet Gesture Distance (FGD) to measure personalization and motion quality. Our approach improves SRA from 27.6\% for the frozen prior to 69.5\% while preserving motion quality (FGD 34.2 versus 34.8), and replacing the enrollment embedding with another person's reduces SRA to 11.4\%. The generated gestures are retargeted to a physical NAO robot, and this improvement also transfers to the robot deployment setting, where speech is synthesized using five TTS voices, retaining 67.3\% SRA.