日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.10506

DUET-DINO: ロボットマニピュレーションの潜在計画のための同時クロスビュー世界モデリング

DUET-DINO: Simultaneous Cross-View World Modeling for Latent Planning in Robot Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

側面カメラと手首カメラの観測をクロスビュー条件付けで統合し、7自由度エンドエフェクタ制御のための行動条件付き潜在世界モデルを学習する手法を提案。到達・角度付き到達・把持持ち上げタスクで単一視点や独立二視点のベースラインを上回る性能を達成した。

詳しい要約

1. どんなもの?

- 7-DoF end-effector 制御のための action-conditioned latent world model - 静的な side-camera と wrist-camera の観測から cross-view conditioning で同時に予測を学習 - 全 7-DoF 行動空間での latent planning を可能にする DUET-DINO を提案 - reach, angled-reach, grasp-and-lift タスクで評価

2. 先行研究と比べてどこがすごい?

- 従来の action-conditioned latent world model は fine-grained な空間・回転行動の予測が不正確で 7-DoF 制御に不向き - single-view および independent dual-view baseline を一貫して上回る - 成功率 92% (reach), 72.5% (angled-reach), 60.0% (lift) - V-JEPA 2 の wrist-view 予測が fine-grained 行動による視覚変化を過小評価するのに対し、DINOv3 予測がより良く捉えることを示す

3. 技術・手法の肝は?

- static side-camera と wrist-camera の観測を cross-view conditioning で統合 - グローバルなシーン情報と gripper 中心の情報を相補的に活用 - action-conditioned な未来の視覚表現を同時に予測する latent world model - DROID と RoboArena データセットでゼロから学習

4. どうやって有効だと検証した?

- 空間的に多様な reach、回転重視の angled-reach、multi-goal の grasp-and-lift タスクで評価 - single-view および independent dual-view baseline と比較 - 成功率 92% (reach), 72.5% (angled-reach), 60.0% (lift) を達成 - 視覚的分布シフト下での頑健な汎化を確認

5. 議論はある?

- V-JEPA 2 の wrist-view 予測は fine-grained 行動による視覚変化を過小評価 - DINOv3 予測は action-conditioned なシーン変化をより良く捉え、下流の planning を強化 - コードとモデルチェックポイントはオープンソース化予定 - その他の議論や限界は要旨からは不明

6. 次に読むべき論文は?

- V-JEPA 2 - DINOv3 - DROID - RoboArena - action-conditioned latent world model に関する関連研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Nisarga Nilavadi, Ralf Römer, Moritz Reuss, Michael Krawez, Tobias Jülg, Angela P. Schoellig, Rudolf Lioutikov, Wolfram Burgard

分類: cs.RO, cs.CV

原文アブストラクト

Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However, their predictions for fine-grained spatial and rotational actions are unreliable for full 7-DoF end-effector control. To address this gap, we introduce DUET-DINO, a simultaneous cross-view latent world model that jointly learns action-conditioned predictions from static side- and wrist-camera observations through cross-view conditioning. By exploiting complementary global scene and gripper-centric information, DUET-DINO enables latent planning over the full 7-DoF action space. Across spatially diverse reach, orientation-intensive angled-reach, and multi-goal grasp-and-lift tasks, DUET-DINO consistently outperforms single-view and independent dual-view baselines, achieving 92% success on reach, 72.5% on angled-reach, and 60.0% on lift tasks. DUET-DINO is trained from scratch on DROID and RoboArena datasets and generalizes robustly under visual distribution shifts. We further show that while V-JEPA 2 wrist-view predictions underestimate visual dynamics induced by fine-grained actions, DINOv3 predictions better capture action-conditioned scene changes, leading to stronger downstream planning. The code and model checkpoints will be open-sourced. Project page: https://utn-air.github.io/DUET-DINO

関連論文