日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
模倣学習/人間からロボットへの転移arXiv:2609.21514

Skel-WAM: 人間からロボットへの操作スキル転移を実現する手骨格条件付きワールドアクションモデル

Skel-WAM: A Hand-Skeleton-Conditioned World Action Model for Human-to-Robot Manipulation Transfer

シェア:XThreadsFacebookLINEはてブBluesky

人間とロボットの手の骨格を共通インターフェースとして用い、人間の動画とロボットの実演を統合学習することで、実機の双腕操作タスクの成功率を大幅に向上させた研究。

詳しい要約

1. どんなもの?

- 人間動画とロボット実演を統合学習する World Action Model『Skel-WAM』 - 手の骨格運動を共通インターフェースとして人間とロボットの動作を整列 - 視覚・骨格動態を Video/Keypoint Experts が Mixture-of-Transformers で学習 - 別途 robot-trained Action Expert が予測を実行可能な制御に変換 - 4つの実世界両手タスクと7つのシミュレーションタスクで評価

2. 先行研究と比べてどこがすごい?

- 人間動画は低コストだが visual appearance と action space の embodiment gap が障壁 - 従来は人間動画にロボット action label が必要、または分布カバレッジが限定的 - Skel-WAM は共通 hand topology で人間とロボットの運動を整列し、人間動画に action label 不要 - 実世界で最強 baseline を 22.22 ポイント、シミュレーションで 8.28 ポイント上回る - 人間-ロボット cotraining でロボット訓練に無いタスク変種の成功率が 38.89%→86.11% に倍増

3. 技術・手法の肝は?

- 共通 hand topology による hand-skeleton motion interface を導入 - skeleton overlays で運動をシーンに接地、structured 2.5-D keypoints で手 kinematics を明示 - Video Expert と Keypoint Expert が Mixture-of-Transformers で視覚・骨格動態を共同学習 - 別の robot-trained Action Expert が予測を実行可能な制御にマッピング - この分離により人間動画とロボット実演が共有 dynamics を直接監督可能

4. どうやって有効だと検証した?

- 4つの実世界両手タスクと7つのシミュレーションタスクで評価 - 平均成功率 79.86%(実世界)と 63.29%(シミュレーション)を達成 - 最強 baseline をそれぞれ 22.22 ポイント、8.28 ポイント上回る - 人間-ロボット cotraining でロボット訓練に無いタスク変種の成功率が 38.89%→86.11% に向上

5. 議論はある?

- 共有 skeletal interface が人間とロボットデータの joint learning を可能にすると主張 - 人間実演の補完によりロボットのタスクカバレッジを拡大 - 限界・失敗事例・計算コスト・スケーラビリティに関する議論は要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない - 関連手法として Mixture-of-Transformers、World Action Model、human-to-robot transfer、cotraining の定番研究を挙げる - 具体的な論文名は要旨からは不明

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zetao Cai, Yaping Li, Yiqun Wang, Xinyu Zhan, Yuyin Yang, Haoxiang Ma, Kailin Li, Tao Lu, Jiangmiao Pang, Linning Xu, Dahua Lin

分類: cs.RO

原文アブストラクト

Robot demonstrations are expensive to collect and often provide limited distributional coverage of task variations. Human videos offer a low-cost source of complementary manipulation experience, but learning from them requires bridging embodiment gaps in visual appearance and action spaces. We introduce Skel-WAM, a world action model that bridges these differences through a unified hand-skeleton motion interface. The key insight is to align human and robot motion through a common hand topology, combining skeleton overlays that ground motion in the scene with structured 2.5-D keypoints that encode explicit hand kinematics. Video and Keypoint Experts jointly learn visual and skeletal dynamics through a Mixture-of-Transformers, while a separate robot-trained Action Expert maps these predictions to executable controls. This separation enables human and robot demonstrations to directly supervise shared dynamics without requiring robot action labels for human videos. Across four real-world bimanual tasks and seven simulated tasks, Skel-WAM achieves average success rates of 79.86% and 63.29%, surpassing the strongest baseline by 22.22 and 8.28 percentage points, respectively. Human-robot cotraining more than doubles real-world success on task variations absent from robot training data, from 38.89% to 86.11%. These results demonstrate that a shared skeletal interface enables joint learning across human and robot data and expands robot task coverage through complementary human demonstrations.

PR本紙発行元 EmplifAI