日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.20396

Imagine-TAMP:部分観測下での想像誘導によるタスク・動作計画

Imagine-TAMP: Imagination-Guided Task and Motion Planning in Partial Observability

シェア:XThreadsFacebookLINEはてブBluesky

視覚言語モデルと生成シーンモデルで未観測領域を想像し、観測と操作のどちらを優先すべきかを計画段階で比較するTAMPフレームワークを提案。

詳しい要約

1. どんなもの?

部分観測下のロボット操作のための、想像に基づくTask and Motion Planning (TAMP) フレームワーク。 - 対象物の位置が部分的にしか観測できない環境で、追加観察か遮蔽物操作かを判断する。 - 視覚言語モデルと生成シーンモデルによる意味的・幾何的想像を用いる。 - 計画と実行を交互に行い、信念更新と再計画を統合する。

2. 先行研究と比べてどこがすごい?

従来のTAMPは象徴的action costや高価な幾何計画で観察か操作かを決めていた。 - それらは観察が遮蔽対象を明らかにする確率を適切に捉えない。 - Imagine-TAMPは想像に基づく非単位コストで戦略を比較し、高価なmotion planning前に判断する。 - 観察と操作のトレードオフをより良く扱う点が新しい。

3. 技術・手法の肝は?

視覚言語モデルが対象と可視物体の常識関係を用いて対象位置のparticle beliefを形成する。 - 生成シーンモデルが未観測領域の妥当な幾何を推定する。 - 対象仮説と想像シーンから複数の象徴的plan skeletonを生成する。 - 操作労力と観察による対象可視性を近似する非単位コストを割り当てる。 - 選択したskeletonを連続計画に洗練し実行、新観測で信念更新と再計画を行う。

4. どうやって有効だと検証した?

視点制約のあるshelfシーンで実験。 - 非単位幾何評価により成功率が46.0%から84.0%に向上。 - 意味的信念形成が操作と再計画をさらに削減。 - 実ロボットで完全システムがgeometry-only ablation比で計画時間を32%削減。

5. 議論はある?

要旨からは不明。 - 限界や失敗事例、計算コストの詳細な議論は記述されていない。

6. 次に読むべき論文は?

要旨で参照/比較されている研究は明示されていない。 - 関連手法としてTask and Motion Planning (TAMP)、vision-language model、generative scene model、particle belief、geometry-only ablationが挙げられる。 - 同分野の定番として部分観測下のTAMPや能動的知覚に関する研究を読むと良い。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Antareep Singha, Shivaram Kumar, Yoonwoo Kim, Yoonchang Sung

分類: cs.RO

原文アブストラクト

Robots operating in cluttered environments must often manipulate objects whose locations are only partially observable. A central challenge is deciding whether to acquire another observation or to first manipulate objects that may occlude the target. Conventional task and motion planning (TAMP) approaches typically make this decision using symbolic action costs or expensive geometric planning, neither of which adequately captures how likely an observation is to reveal an occluded target. We introduce Imagine-TAMP, an interleaved planning and execution framework that uses semantic and geometric imagination to compare alternative task-level strategies under partial observability before committing to expensive motion planning. A vision-language model shapes a particle belief over target locations using commonsense relationships between the target and visible objects, while a generative scene model estimates plausible geometry in unobserved regions. Given a target hypothesis and imagined scene, Imagine-TAMP generates multiple symbolic plan skeletons and assigns non-unit costs that approximate both manipulation effort and target visibility from sensing actions, distinguishing a short but poorly informative observation strategy from a longer strategy that first manipulates an occluder to better expose the target. The selected skeleton is then refined into a feasible continuous plan and executed, with new observations updating the belief and triggering replanning when necessary. Experiments show that imagination-guided evaluation improves observation-versus-manipulation decisions: in viewpoint-constrained shelf scenes, non-unit geometric evaluation increases success from 46.0% to 84.0%, while semantic belief shaping further reduces manipulation and replanning. On a real robot, the complete system reduces planning time by 32% relative to a geometry-only ablation.

関連論文

PR本紙発行元 EmplifAI