日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
動作生成arXiv:2608.28213v1

PAMoR: ヒューマノイドロボットのためのリアルタイムパラメトリック感情動作生成

PAMoR: Parameterized Affective Motion Generation in Real Time for Humanoid Robots

シェア:XThreadsFacebookLINEはてブBluesky

ロボットの動作に感情を定量的にパラメータ化して組み込む手法を提案。価覚醒座標を制御パラメータとして用い、リアルタイムで全身動作を生成する。

詳しい要約

1. どんなもの?

PAMoRは、ヒューマノイドロボットの動作生成において、感情を定量的な制御パラメータとして扱う手法を提案する。具体的には、valence-arousal (V-A) 座標をロボットの運動学から閉形式で計算し、これを生成条件として用いる。動作生成は、29-DoFのUnitree G1上でリアルタイムに自己回帰的に行われ、動作と感情の両方を編集可能である。

2. 先行研究と比べてどこがすごい?

従来の感情動作生成は、人間のアバター向けであり、スタイルを参照クリップや感情語から取得していたため、定量的なパラメータ化ができなかった。PAMoRは、V-A座標を直接ロボットの運動学から計算し、人手によるアノテーションなしで生成条件とする点が新しい。また、リアルタイムで全身動作を生成できる点も先行研究と異なる。

3. 技術・手法の肝は?

手法の核は、共有潜在空間で訓練されたアクション事前分布と2つの感情事前分布を、各denoisingステップで合成することである。アクション事前分布は「何をするか」を固定し、感情事前分布は「どのようにするか」を変調する。V-A座標は、姿勢の拡がりと運動エネルギーから閉形式で計算され、生成条件として直接使用される。

4. どうやって有効だと検証した?

有効性は、生成された動作が指示されたV-A座標を全範囲にわたって追跡すること、およびtext-to-motionの忠実度がテキストのみのベースラインと同等であることで検証された。さらに、知覚実験では、評価者が指示された感情を0.38の試行で識別し、これはベースラインを上回り、実際の人間の身体動作で報告された0.44に近い値である。

5. 議論はある?

要旨からは、議論の余地や限界についての詳細は不明である。ただし、知覚実験での識別率が人間の動作より低いことから、感情表現の自然さや明瞭さには改善の余地がある可能性が示唆される。また、V-A座標の計算がロボットの運動学に依存するため、他のロボットへの一般化については要旨からは不明である。

6. 次に読むべき論文は?

要旨で参照されている研究は明示されていないが、関連する分野として、感情動作生成のためのtext-to-motionモデルや、人間の動作からの感情認識に関する研究が挙げられる。具体的には、HumanML3DやAIST++などのデータセットを用いた研究や、動作生成における拡散モデルの応用が関連する。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yan Pan, Lingfan Bao, Tianhu Peng, Chengxu Zhou

分類: cs.RO

原文アブストラクト

People read a humanoid robot's motion in social settings not only for the action performed but for the affect conveyed. Motion carrying that affect has so far been generated for human avatars, where style is taken from a reference clip or an emotion word, neither of which can be quantitatively parameterized. We present PAMoR, which turns affect into a measured control parameter: a valence-arousal (V-A) coordinate computed natively on robot kinematics. It is obtained in closed form from postural expansion and movement energy, and these measurements serve directly as generation conditions, with no human annotation. An action prior and two affect priors, trained in a shared latent space, are composed at each denoising step: the action prior fixes what is performed, the affect priors modulate how. Whole-body motion rolls out autoregressively on a 29-DoF Unitree G1 in real time, with action and affect both editable. Generated motion tracks the commanded V-A over its full range while text-to-motion fidelity still matches text-only baselines. In a perceptual study, raters identify the commanded emotion on 0.38 of trials, above both baselines and approaching the 0.44 reported for acted human bodies.