日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
計画arXiv:2607.08024v1

APIVOT: 視覚と言語の思考を適応的に交互配置する長期的計画手法

APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts

シェア:XThreadsFacebookLINEはてブBluesky

ロボットの長期的なタスク計画において、言語による意味的推論と視覚による幾何学的検証を適応的に交互に行うVLMベースのプランナーを提案し、キッチンタスクで性能向上を実証した。

著者: Emily Jin, Joy Hsu, Yiqing Xu, Weiyu Liu, Nick Haber, Jiajun Wu

分類: cs.CV, cs.AI, cs.LG, cs.RO

原文アブストラクト

Long-horizon robot planning requires jointly reasoning over semantic task structure and geometric feasibility. To successfully execute a task, a robot must decompose goals, select task-relevant objects, and sequence actions, while ensuring that plans satisfy spatial constraints such as limited free space and object collisions. In this work, we propose APIVOT, a VLM-based planner that adaptively interleaves language and visual thoughts for long-horizon planning. APIVOT learns to leverage language for semantic reasoning, while using visual thoughts as imagined future states for internal verification of geometric feasibility. On long-horizon kitchen tasks, APIVOT outperforms general-purpose VLMs and prior planning frameworks, achieving the largest gains in spatially constrained settings. We find that APIVOT learns meaningful modality selection behavior, demonstrating that adaptive interleaving of vision-language thoughts improves both planning success and reasoning efficiency.

関連論文