日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ヒューマノイド/ジェスチャー生成arXiv:2609.23414

EmoPose: 視覚言語モデルが導く感情を考慮したヒューマノイドロボットのジェスチャー生成

EmoPose: Vision-Language Model Guided Emotion-Aware Gesture Generation for Humanoid Robots

シェア:XThreadsFacebookLINEはてブBluesky

視覚言語モデルがジェスチャーの種類や強度を選び、ロボット側の動作ライブラリで実行する枠組みを提案し、社会的文脈に応じた表情豊かなジェスチャーを実現した。

著者: Daojie Peng, Bingtao Wang, Fulong Ma, Wenjun Yue, Liang Zhang, Jun Ma

分類: cs.RO

原文アブストラクト

Socially competent humanoid robots must communicate affect and intent through gesture as well as speech, yet open-ended interaction must become motion that is both expressive and executable on a specific body. This demands semantic flexibility for contextual social intent while preserving deterministic, embodiment-aware robot control. We present EmoPose, a vision-language model (VLM)-guided framework that bridges this gap through an executable semantic interface. Given language, dialogue history, and optional visual context, the VLM selects an ordered gesture plan containing a communicative class, library variant, intensity, and speech anchor. A scalable robot-owned motion library defines the available expressive vocabulary and the source of 14-DoF joint targets. Pose Studio supports automatic trajectory generation, MuJoCo preview, and automatic synchronization of new library entries with the VLM guide; deterministic robot-side modules validate plans, construct trajectories, schedule gestures, and manage queueing and interruption. This division lets the interaction repertoire grow for new social contexts without changing the control interface or delegating raw joint commands to the foundation model. On the EmoPose-Bench, structured GPT-5.5 planning reaches $98.25\pm0.52\%$ on the Easy tier and $76.50\pm0.54\%$ overall, exceeding same-model direct-label prompting. Further tests validate dialogue-context use and ordered multi-action composition. The system completes the nominal MuJoCo suite and realizes all 29 authored variants on the physical Unitree G1. A four-stop laboratory tour demonstrates expressive narration with interruption, camera-grounded dialogue, and navigation.

PR本紙発行元 EmplifAI