JAMB:両手操作のための動作と運動の同時拡散
JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
両手の動作と将来の3D点追跡を共有Transformer内で同時にノイズ除去し、相互に情報を活用して両手操作の精度と汎化性能を高める拡散ポリシーを提案。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Chuyang Xiao, Peilin Meng, David Held
分類: cs.RO
原文アブストラクト
Coordinated bimanual manipulation is challenging because the motion of either arm can alter the shared 3D scene and thereby affect the other arm. Yet most diffusion policies generate actions without explicitly modeling these future geometric consequences, while predictive variants typically use future state only as auxiliary supervision or fixed conditioning. We address this limitation by proposing JAMB, a diffusion policy that jointly denoises bimanual actions and future 3D point tracks. By allowing action and track hypotheses to evolve together within a shared Transformer, each can inform and refine the other throughout denoising. We further ground multimodal representations in a shared spatiotemporal coordinate system to facilitate geometry-aware interaction during joint denoising. We evaluate JAMB on diverse bimanual manipulation tasks in RoboTwin 2.0 and on a real-world robot, comparing it with action-only policies and alternative future-prediction approaches spanning different state representations and learning objectives. Across 16 simulation tasks, JAMB achieves an average success rate of 83.4%, outperforming the strongest baseline by 23.9 percentage points. On three real-world tasks, it outperforms the action-only and auxiliary geometry prediction methods by 50.0 and 21.2 percentage points, respectively. Beyond these performance gains, JAMB shows stronger generalization to cluttered scenes and out-of-distribution backgrounds than the evaluated baselines. Together, these results demonstrate the effectiveness of our joint action-motion modeling framework for coordinated bimanual manipulation. Our project website is available at https://jam-bimanual.github.io/
関連論文
- エンドタスク成功を超えて:ロボティクスにおける視覚経験検索の監査手法マニピュレーション
- 複数把持点における非無視可能な物理応答を伴う線形変形物体の安定性保証付きマニピュレーションマニピュレーション
- PAKT: 強化学習のための物理的整合性を備えたキネステティック教示マニピュレーション
- 不完全データを活用した高精度ロボットマニピュレーションマニピュレーション
- SafeLoop: 視覚言語行動マニピュレーションのためのリスク認識ロールバックマニピュレーション
- ローカルコーディングエージェントによるマニピュレーションスキルの汎化マニピュレーション