日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.31112

DualManip:二重経路の意味推論と幾何適応によるエージェント型動的マニピュレーション

DualManip: Agentic Dynamic Manipulation via Dual-Path Semantic Reasoning and Geometric Adaptation

シェア:XThreadsFacebookLINEはてブBluesky

VLMによる意味推論と形状適応ネットワークによる幾何適応を分離し、動的な物体変化に追従しながら把持を再構成するロボットマニピュレーション手法を提案。

詳しい要約

1. どんなもの?

- VLMによるopen-vocabularyなロボットmanipulationフレームワーク。 - 意味推論と幾何適応を分離したdual-path構成。 - 動的sceneでの応答性とrobustnessを両立。 - 非剛体変形・関節再構成・剛体運動・高精度assemblyを対象。

2. 先行研究と比べてどこがすごい?

- VLMの高い推論遅延が動的sceneでの応答性を制限する課題に対処。 - 頻繁でないsemantic reasoningと応答的なgeometric adaptationを分離。 - 連続的なscene変化下で優れたmanipulation robustnessを実証。 - geometric adaptationはagentic verificationとsemantic replanningより約46倍高速。

3. 技術・手法の肝は?

- semantic path: タスク分解とtask-relevant interactionのgrounding、constraint-solvingによるpose最適化。 - geometric path: 生のRGB-D観測からshape-adaptive networkでtemplate-to-observation対応を継続更新。 - 対応関係を介してtask-relevant grasp contactsを転送し、物体運動・非剛体変形下でonline grasp reconstruction。 - Information Interaction Moduleが両pathを橋渡し: semantic groundingからgrasp初期化、geometric更新の検証、更新失敗時のsemantic replanning起動。

4. どうやって有効だと検証した?

- 実世界評価で6つのmanipulationタスクを実施。 - 非剛体変形、articulated reconfiguration、剛体運動、高精度assemblyをカバー。 - static、single-change、continuous dynamicの3設定で評価。 - 連続的scene変化下で優れたrobustnessを示し、geometric adaptationが約46倍高速であることを確認。

5. 議論はある?

- 要旨からは不明。 - 限界や失敗事例、計算コストの詳細な議論は記述されていない。

6. 次に読むべき論文は?

- 要旨で参照・比較されている研究は明示されていない。 - 関連手法としてVLMベースのmanipulation、open-vocabulary reasoning、agentic verification、semantic replanning、shape-adaptive network、RGB-D対応推定などが挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chengxi Li, Yan Di, Yingyue Li, Ruida Zhang, Mingyang Li, Xiangyang Ji

分類: cs.RO

原文アブストラクト

Vision-language models (VLMs) enable open-vocabulary reasoning for robot manipulation, but their high inference latency limits responsiveness in dynamic scenes. Many scene changes, however, alter object geometry without invalidating task intent. We present DualManip, a dual-path framework that decouples infrequent semantic reasoning from responsive geometric adaptation. The semantic path decomposes the task and grounds task-relevant interactions, followed by a constraint-solving module for pose optimization. During execution, the geometric path continuously updates template-to-observation correspondences from live RGB-D observations via a shape-adaptive network. These correspondences transfer task-relevant grasp contacts across observations, enabling online grasp reconstruction under object motion and non-rigid deformation. The Information Interaction Module bridges the two paths by initializing task-relevant grasps from semantic grounding, validating geometric updates, and triggering semantic replanning upon update failures. Real-world evaluation spans six manipulation tasks covering non-rigid deformation, articulated reconfiguration, rigid motion, and high-precision assembly across three settings: static, single-change, and continuous dynamic. DualManip demonstrates superior manipulation robustness, particularly under continuous scene changes, while achieving geometric adaptation approximately 46$\times$ faster than agentic verification and semantic replanning. Our project page: https://lichengxi1.github.io/Dualmanip.

関連論文

PR本紙発行元 EmplifAI