SkeleWAM: 骨格ワールドアクションモデリングによる効率的なロボットマニピュレーション
SkeleWAM: Skeleton World-Action Modeling for Efficient Robotic Manipulation
ロボット関節・物体中心・接触点からなる疎な3D骨格を状態表現として用い、行動生成と将来骨格予測を統合した軽量なワールドアクションモデルを提案。LIBERO-Plusで85.9%の成功率を達成。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Juyi Sheng, Hua Wang, Mengyuan Liu
分類: cs.RO
原文アブストラクト
World action models (WAMs) combine robot action generation with future state prediction. Existing WAMs typically predict videos or learned visual latents, which represent interaction geometry only implicitly and may retain appearance information unrelated to control. We introduce SkeleWAM, a compact WAM that represents a manipulation scene as a sparse 3D skeleton composed of robot joints, object centers, and interaction points. Constructed online from current RGB-D observations and robot proprioception, the skeleton provides a unified geometric state for action generation and future skeleton prediction. Future skeleton prediction provides additional geometric supervision for action learning without requiring visual reconstruction. At inference, SkeleWAM generates actions directly from the current skeleton and language instruction, while Medoid Action Consensus (MAC) serves as an auxiliary consensus strategy for stochastic action samples. On LIBERO-Plus, SkeleWAM achieves an overall success rate of 85.9% with 57.1M parameters, outperforming Cosmos-Policy by 3.7 percentage points. These results demonstrate that sparse 3D robot--object structure provides an effective state space for robust and parameter-efficient world action learning. The project is available at https://skelewam-project.github.io/.
関連論文
- スクリューアテンション:Transformer内部に剛体代数を組み込むマニピュレーション
- FlashDexRetarget: 多動作リターゲティングによる器用操作データ生成の高速化マニピュレーション
- 解像度に一貫したヤコビアン場を学習する生体模倣剛柔指マニピュレーション
- Recova: 自律ロボットマニピュレーションのためのエージェント誘導型失敗回復マニピュレーション
- 経験と実演による6自由度把持合成の継続学習マニピュレーション
- 再構成・練習・実世界展開:身体性エージェントのためのガイド付き自己改善マニピュレーション