日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
動作生成/リターゲティングarXiv:2609.37297

クロススケルトン動作リターゲティングが非識別可能である理由:生成動作モデルの構造的限界

Why Cross-Skeleton Retargeting Is Non-Identifiable: Structural Limits of Generative Motion Models

シェア:XThreadsFacebookLINEはてブBluesky

異なる骨格間で動作を転送する生成モデルは、ソース動作を反映したのか典型的動作を生成したのか区別できない構造的曖昧性を持つことを示し、その検証指標SIFを提案した。

詳しい要約

1. どんなもの?

異なる骨格間で動作を転送する cross-skeleton retargeting の生成モデルを扱う研究。source clip の動作構造と意図を別の身体へ移す際、出力が「source を転送した」のか「要求された action の典型動作を生成した」のかを訓練データから区別できないという構造的曖昧性を示す。 - 対象: cross-skeleton motion generation - 主張: source-conditioned retargeting map は sparse heterogeneous motion domains で non-identifiable - 提案: Source-Instance Fidelity (SIF) という診断指標

2. 先行研究と比べてどこがすごい?

先行研究との具体的比較は要旨からは不明。 - 従来は action-level のテスト(正しい action か)で評価されることが多いと示唆 - 本論文は、その評価では source 依存性を捉えられないと指摘 - 曖昧性を偶発的でなく構造的(non-identifiable)と定式化した点が特徴

3. 技術・手法の肝は?

生成目的下での identifiability を理論的に分析し、診断指標を導入する。 - Unpaired distribution matching: 異なる骨格の latent space を相対変換しても training evidence が変わらず、gauge non-identifiability が生じる - Sparse paired supervision: action のみで pair すると squared-error 訓練が source clip を無視する平均動作へ収束(conditional-mean degeneration) - SIF: target skeleton と action を固定し、出力間の差異が source clip 間の差異と対応するかを検査

4. どうやって有効だと検証した?

animal motion data 上で SIF により検証。 - 標準の action-level テストに成功する手法の多くが source-blind floor に留まる - source-blind floor を上回る手法でも relational signal を部分的にしか保持しない - 具体的なデータセット名・ベースライン名・数値は要旨からは不明

5. 議論はある?

retargeting には、主張する source-conditioned map を同定できる目的関数と評価が必要だと議論。 - 現行の action-level 評価の限界 - unpaired と sparse paired の相補的な失敗モード - 具体的な限界・反論・倫理的議論は要旨からは不明

6. 次に読むべき論文は?

要旨で参照・比較されている個別研究は明示されていない。 - cross-skeleton retargeting / motion retargeting の生成モデル - unpaired distribution matching を用いる motion generation - identifiability や causal representation learning の関連研究 - 同分野の定番として human motion retargeting や generative motion models の文献

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhiyuan Li, Wenyan Yang, Pekka Marttinen, Joni Pajarinen

分類: cs.CV, cs.RO

原文アブストラクト

Cross-skeleton motion generation trains generative models to carry action structure and motion intention from one body to another. Yet a target motion that shows the right action has two explanations that the training data cannot tell apart: the model transferred the source clip, or it recovered a typical motion for the requested action. We show that this ambiguity is structural rather than incidental: under standard generative objectives, the source-conditioned retargeting map is non-identifiable in sparse heterogeneous motion domains. Unpaired distribution matching yields gauge non-identifiability: the latent spaces of different skeletons can be transformed relative to one another without changing the training evidence, so different source-conditioned maps fit it equally well. Sparse paired supervision admits the complementary failure mode, \emph{conditional-mean degeneration}: when clips are paired only by action, squared-error training converges to an average target motion that ignores the source clip. To make the missing evidence observable, we introduce Source-Instance Fidelity (SIF), a diagnostic that tests whether outputs differ from one another the way their source clips do, with the target skeleton and action held fixed. Under this diagnostic, methods that succeed at the standard action-level test on animal motion data often sit at the source-blind floor, while the methods that rise above it retain only a partial relational signal. Retargeting therefore needs objectives and evaluations that can identify the source-conditioned map it claims to learn. Project page: https://cross-skeleton-retargeting.netlify.app/.

PR本紙発行元 EmplifAI