日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
異形態間転移arXiv:2609.21983

SkelWAM: 骨格ガイド型ワールドアクションモデルによるゼロショット異形態間マニピュレーション

SkelWAM: A Skeleton-Guided World-Action Model for Zero-Shot Cross-Embodiment Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

腕中心線・TCP姿勢・把持コマンドを共有25次元状態として用い、映像と行動のTransformerで骨格行動を予測し、形態別デコーダで異なるロボットの制御に変換するゼロショット転移手法を提案。

詳しい要約

1. どんなもの?

SkelWAMは、単一のソースembodimentで学習した操作経験を、目標タスクのデモンストレーションやポリシー更新なしに他embodimentへゼロショット転移する、skeleton-guided world-action modelである。 - 知覚と制御を1つの明示的幾何表現で結合する - arm centerline geometry、tool-center-point (TCP) pose、parallel-jaw commandsからなる共有25-D stateを用いる - 同一定義がcanonical third-person観測、wrist観測、将来のwhole-body action targetsの基礎となる - 予測的視覚supervisionで学習したvideo-action mixture of transformersがcanonical skeleton action chunksを予測する - embodiment固有のconstrained decoderがjointまたはcontinuum-robot制御へ変換する - 一対一のjoint対応を必要…

2. 先行研究と比べてどこがすごい?

要旨では、embodimentが変わると視覚的外観、action次元と意味、同じtool poseを実現するwhole-body configurationが変化する点が課題として挙げられている。 - SkelWAMは単一ソースのcross-embodiment manipulationを対象とする - 一対一のjoint correspondenceを必要としない点を特徴とする - 目標タスクのデモンストレーションや目標ポリシー更新を不要とする - 新benchmark LIBERO-Cross10上で、Franka学習済みSkelWAMが1,000エピソードで43.3%成功し、評価した最良baselineを36.2 percentage points上回る - 具体的な先行研究名との比較は要旨からは不明

3. 技術・手法の肝は?

技術の肝は、知覚と制御を1つの明示的幾何表現で結ぶskeleton-guided world-action modelにある。 - arm centerline geometry、TCP pose、parallel-jaw commandsを共有25-D stateとして定義する - この定義をcanonical third-person観測、wrist観測、将来のwhole-body action targetsに共通して用いる - 予測的視覚supervisionで学習したvideo-action mixture of transformersがcanonical skeleton action chunksを予測する - embodiment固有のconstrained decoderが、そのchunksをjointまたはcontinuum-robot制御へ変換する - 一対一のjoint対応を不要とし、目標タスクのデモや目標ポリシー更新を用いない

4. どうやって有効だと検証した?

検証は新benchmarkと実機展開で行っている。 - LIBERO-Cross10を導入する。これはsource-only cross-embodiment transfer benchmarkで、10タスクと4 morphological groupsにわたる10 target embodimentsを対象とする - このbenchmark上で、Franka学習済みSkelWAMが1,000エピソードで43.3%成功し、評価した最良baselineを36.2 percentage points上回った - さらにJAKA mini2学習済みポリシーをFeagine A03 continuum robotへ展開し、3つのtabletop manipulationタスクで実世界cross-embodiment manipulationの可能性を示した

5. 議論はある?

要旨では、SkelWAMが一対一のjoint対応を必要とせず、目標タスクのデモや目標ポリシー更新なしに単一ソースからcross-embodiment転移できる点を主張している。 - LIBERO-Cross10上で43.3%成功、最良baselineを36.2 percentage points上回る結果を示す - JAKA mini2からFeagine A03 continuum robotへの実機展開で可能性を示す - 限界、失敗要因、制約、倫理的・実運用上の議論は要旨からは不明

6. 次に読むべき論文は?

要旨で参照・比較されている具体的な先行研究名は明示されていない。 - 関連手法として、cross-embodiment manipulation、world-action model、video-action mixture of transformers、skeleton-guided representation、LIBERO benchmark系の研究を挙げる - 特にLIBERO-Cross10やsource-only cross-embodiment transfer benchmarkに関連する研究が次の読むべき候補となる - 具体的な論文名は要旨からは不明

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Pengjun Niu, Yujia Xie, Rui Peng, Hang Zhao, Ke Liu

分類: cs.RO

原文アブストラクト

Reusing manipulation experience across robot embodiments is important for scaling robot learning and reducing repeated task-specific data collection. However, changes in embodiment alter visual appearance, action dimensionality and semantics, and the whole-body configurations that can realize the same tool pose. We present SkelWAM, a skeleton-guided world-action model that couples perception and control through one explicit geometric representation for single-source cross-embodiment manipulation. Arm centerline geometry, tool-center-point (TCP) pose, and parallel-jaw commands form a shared 25-D state. The same definition underlies canonical third-person and wrist observations and future whole-body action targets. Trained with predictive visual supervision, a video-action mixture of transformers predicts canonical skeleton action chunks, which embodiment-specific constrained decoders convert into joint or continuum-robot controls. This formulation requires no one-to-one joint correspondence and uses no target-task demonstrations or target policy updates. We introduce LIBERO-Cross10, a source-only cross-embodiment transfer benchmark covering ten tasks and ten target embodiments across four morphological groups. On this benchmark, Franka-trained SkelWAM achieves 43.3% success over 1,000 episodes, exceeding the best-performing evaluated baseline by 36.2 percentage points. We further deploy a JAKA mini2-trained policy on the Feagine A03 continuum robot for three tabletop manipulation tasks, illustrating the approach's potential for real-world cross-embodiment manipulation. Project page: http://www.liukepku.com/skelwam/index.html

PR本紙発行元 EmplifAI