日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.25322

JAMB:両手操作のための動作と運動の同時拡散

JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

両手の動作と将来の3D点追跡を共有Transformer内で同時にノイズ除去し、相互に情報を活用して両手操作の精度と汎化性能を高める拡散ポリシーを提案。

詳しい要約

1. どんなもの?

JAMBは、双腕マニピュレーションのための拡散ポリシー。双腕のactionと将来の3D point tracksを共有Transformer内でjointly denoiseする。各腕の動きが共有3D sceneを変え他方に影響する問題に対処。RoboTwin 2.0と実機で評価。

2. 先行研究と比べてどこがすごい?

- 多くのdiffusion policyは将来の幾何的影響を明示せずactionのみ生成 - 予測型変種はfuture stateを補助監督や固定conditioningに使うのみ - JAMBはactionとtrack仮説をdenoising中に相互に洗練 - 16 simタスクで平均成功率83.4%、最強baselineを23.9pt上回る - 実機3タスクでaction-onlyとauxiliary geometry predictionを各50.0pt、21.2pt上回る - 混雑sceneやOOD背景への汎化も強い

3. 技術・手法の肝は?

- 双腕actionと将来3D point tracksをjointly denoiseするdiffusion policy - actionとtrack仮説を共有Transformer内で一緒に発展させ相互に情報伝達 - multimodal表現を共有時空間座標系にgroundingし、joint denoising中のgeometry-aware interactionを促進

4. どうやって有効だと検証した?

- RoboTwin 2.0の多様な双腕タスク16件で評価 - 実世界ロボット3タスクで評価 - action-only policyやfuture-prediction手法(異なるstate表現・学習目的)と比較 - 成功率、混雑scene・OOD背景への汎化を検証

5. 議論はある?

要旨からは不明。性能向上と汎化は示されるが、限界や失敗事例、計算コスト、track予測誤差の影響などの議論は要旨に記載なし。

6. 次に読むべき論文は?

- action-only diffusion policy - future-prediction approaches with different state representations and learning objectives - RoboTwin 2.0 - 同分野の定番: Diffusion Policy, ACT, ALOHA

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chuyang Xiao, Peilin Meng, David Held

分類: cs.RO

原文アブストラクト

Coordinated bimanual manipulation is challenging because the motion of either arm can alter the shared 3D scene and thereby affect the other arm. Yet most diffusion policies generate actions without explicitly modeling these future geometric consequences, while predictive variants typically use future state only as auxiliary supervision or fixed conditioning. We address this limitation by proposing JAMB, a diffusion policy that jointly denoises bimanual actions and future 3D point tracks. By allowing action and track hypotheses to evolve together within a shared Transformer, each can inform and refine the other throughout denoising. We further ground multimodal representations in a shared spatiotemporal coordinate system to facilitate geometry-aware interaction during joint denoising. We evaluate JAMB on diverse bimanual manipulation tasks in RoboTwin 2.0 and on a real-world robot, comparing it with action-only policies and alternative future-prediction approaches spanning different state representations and learning objectives. Across 16 simulation tasks, JAMB achieves an average success rate of 83.4%, outperforming the strongest baseline by 23.9 percentage points. On three real-world tasks, it outperforms the action-only and auxiliary geometry prediction methods by 50.0 and 21.2 percentage points, respectively. Beyond these performance gains, JAMB shows stronger generalization to cluttered scenes and out-of-distribution backgrounds than the evaluated baselines. Together, these results demonstrate the effectiveness of our joint action-motion modeling framework for coordinated bimanual manipulation. Our project website is available at https://jam-bimanual.github.io/

関連論文

PR本紙発行元 EmplifAI