日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.02120

SkeleWAM: 骨格ワールドアクションモデリングによる効率的なロボットマニピュレーション

SkeleWAM: Skeleton World-Action Modeling for Efficient Robotic Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

ロボット関節・物体中心・接触点からなる疎な3D骨格を状態表現として用い、行動生成と将来骨格予測を統合した軽量なワールドアクションモデルを提案。LIBERO-Plusで85.9%の成功率を達成。

詳しい要約

1. どんなもの?

- ロボット操作のための compact な World Action Model (WAM) である SkeleWAM を提案。 - 操作 scene を robot joints, object centers, interaction points からなる sparse 3D skeleton として表現。 - 現在の RGB-D 観測と robot proprioception から online で構築され、action generation と future skeleton prediction の unified geometric state を提供。 - 推論時は current skeleton と language instruction から直接 action を生成。 - Medoid Action Consensus (MAC) を stochastic action samples の補助 consensus 戦略として用いる。

2. 先行研究と比べてどこがすごい?

- 既存 WAM は video や learned visual latents を予測し、interaction geometry を暗黙的にしか表現せず、control に無関係な appearance 情報を保持しうる。 - SkeleWAM は sparse 3D skeleton を state space とし、visual reconstruction を必要とせず future skeleton prediction による geometric supervision を action learning に与える。 - LIBERO-Plus で 57.1M parameters ながら overall success rate 85.9% を達成し、Cosmos-Policy を 3.7 percentage points 上回る。 - これにより sparse 3D robot--object structure が robust かつ parameter-efficient な world action learning の有効な state space である…

3. 技術・手法の肝は?

- manipulation scene を robot joints, object centers, interaction points からなる sparse 3D skeleton として表現。 - skeleton は current RGB-D observations と robot proprioception から online で構築。 - skeleton は action generation と future skeleton prediction の unified geometric state として機能。 - future skeleton prediction が visual reconstruction なしに action learning への追加 geometric supervision を提供。 - 推論時は current skeleton と language instruction から action を直接生成。 - Medoid Action Consensus (MAC) が stochastic action sampl…

4. どうやって有効だと検証した?

- LIBERO-Plus で評価。 - 57.1M parameters で overall success rate 85.9% を達成。 - Cosmos-Policy を 3.7 percentage points 上回る。 - これにより sparse 3D robot--object structure の有効性を実証。

5. 議論はある?

- 既存 WAM の video や learned visual latents 予測は interaction geometry を暗黙的にしか表現せず、control に無関係な appearance 情報を保持しうる点を議論。 - sparse 3D robot--object structure が robust かつ parameter-efficient な world action learning の有効な state space であると主張。 - その他の限界や議論は要旨からは不明。

6. 次に読むべき論文は?

- Cosmos-Policy (要旨で比較されている)。 - World Action Models (WAMs) の既存研究。 - LIBERO-Plus (評価に用いられた benchmark)。 - Medoid Action Consensus (MAC) に関連する consensus 戦略。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Juyi Sheng, Hua Wang, Mengyuan Liu

分類: cs.RO

原文アブストラクト

World action models (WAMs) combine robot action generation with future state prediction. Existing WAMs typically predict videos or learned visual latents, which represent interaction geometry only implicitly and may retain appearance information unrelated to control. We introduce SkeleWAM, a compact WAM that represents a manipulation scene as a sparse 3D skeleton composed of robot joints, object centers, and interaction points. Constructed online from current RGB-D observations and robot proprioception, the skeleton provides a unified geometric state for action generation and future skeleton prediction. Future skeleton prediction provides additional geometric supervision for action learning without requiring visual reconstruction. At inference, SkeleWAM generates actions directly from the current skeleton and language instruction, while Medoid Action Consensus (MAC) serves as an auxiliary consensus strategy for stochastic action samples. On LIBERO-Plus, SkeleWAM achieves an overall success rate of 85.9% with 57.1M parameters, outperforming Cosmos-Policy by 3.7 percentage points. These results demonstrate that sparse 3D robot--object structure provides an effective state space for robust and parameter-efficient world action learning. The project is available at https://skelewam-project.github.io/.

関連論文

PR本紙発行元 EmplifAI