DUET-DINO: ロボットマニピュレーションの潜在計画のための同時クロスビュー世界モデリング
DUET-DINO: Simultaneous Cross-View World Modeling for Latent Planning in Robot Manipulation
側面カメラと手首カメラの観測をクロスビュー条件付けで統合し、7自由度エンドエフェクタ制御のための行動条件付き潜在世界モデルを学習する手法を提案。到達・角度付き到達・把持持ち上げタスクで単一視点や独立二視点のベースラインを上回る性能を達成した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Nisarga Nilavadi, Ralf Römer, Moritz Reuss, Michael Krawez, Tobias Jülg, Angela P. Schoellig, Rudolf Lioutikov, Wolfram Burgard
分類: cs.RO, cs.CV
原文アブストラクト
Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However, their predictions for fine-grained spatial and rotational actions are unreliable for full 7-DoF end-effector control. To address this gap, we introduce DUET-DINO, a simultaneous cross-view latent world model that jointly learns action-conditioned predictions from static side- and wrist-camera observations through cross-view conditioning. By exploiting complementary global scene and gripper-centric information, DUET-DINO enables latent planning over the full 7-DoF action space. Across spatially diverse reach, orientation-intensive angled-reach, and multi-goal grasp-and-lift tasks, DUET-DINO consistently outperforms single-view and independent dual-view baselines, achieving 92% success on reach, 72.5% on angled-reach, and 60.0% on lift tasks. DUET-DINO is trained from scratch on DROID and RoboArena datasets and generalizes robustly under visual distribution shifts. We further show that while V-JEPA 2 wrist-view predictions underestimate visual dynamics induced by fine-grained actions, DINOv3 predictions better capture action-conditioned scene changes, leading to stronger downstream planning. The code and model checkpoints will be open-sourced. Project page: https://utn-air.github.io/DUET-DINO
関連論文
- 片手で二部品を組み立てるイン・ハンド・アセンブリマニピュレーション
- 生成動画プランをシミュレーションで接地し多様な器用操作コントローラを実現マニピュレーション
- GTA-2: 接地されたタスク軸によるロボットマニピュレーションスキル合成のためのマルチVLMフレームワークマニピュレーション
- 複数ロボットによる全身皮膚鏡画像の自動撮像スキャナマニピュレーション
- FOCIポリシー:関係的操作ポリシーのためのオブジェクト中心相互作用に焦点を当てるマニピュレーション
- 3DWay: 3D一貫性ウェイポイントによるロボット操作の一般化マニピュレーション