日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ナビゲーションarXiv:2610.12126

SuperNav: あらゆるシーンであらゆるタスクに対応するエージェント型ナビゲーションシステム

SuperNav: An Agentic Navigation System for Any Task in Any Scene

シェア:XThreadsFacebookLINEはてブBluesky

MLLMをナビゲーション用に微調整せず、汎用能力を保ったままエージェントハーネスとナビゲーションスキル・ツールを組み合わせ、画像上で直接目的地を指定できる視覚点インターフェースで多様なタスクと未知環境のナビゲーションを実現した。

詳しい要約

1. どんなもの?

SuperNavは、未知環境で多様な人間の要求に対応する汎用サービスロボット向けのagentic navigation system。MLLMをナビゲーション用にfine-tuningせず、pretrained MLLMに専用のagent harnessを組み合わせる。harnessはNavigation Skills、物理interaction用のagent-oriented Tools、task-progressとcontext管理を提供。unified visual-point interfaceで画像内に直接目的地を指定し、実行feedbackから決定を修正できる。

2. 先行研究と比べてどこがすごい?

既存手法はMLLMをfine-tuningしてナビゲーションactionを予測するため、ナビゲーション訓練データのカバレッジに依存し、新しい要求や環境へのgeneralizationが制限されうる。SuperNavはMLLMのgeneral-purpose能力を保ち、motion executionをnavigation toolsに委譲する点が異なる。4つのbaselineをinstance-level、multi-object、demand-drivenタスクで上回った。

3. 技術・手法の肝は?

MLLMは要求解釈、scene理解、意思決定に集中し、motion executionはnavigation toolsに委譲する。pretrained MLLMにナビゲーション特化のfine-tuningをせず、専用agent harnessを付与。harnessはNavigation Skills、agent-oriented Tools、task-progressとcontext管理を備える。unified visual-point interfaceにより、モデルは画像内で目的地を直接指定し、実行feedbackから決定を修正できる。

4. どうやって有効だと検証した?

instance-level、multi-object、demand-drivenタスクで4つの評価baselineを上回った。HM3Dでのcategory-level評価、および実機quadruped robotへのdeploymentにより、環境をまたぐ適用可能性を示した。

5. 議論はある?

要旨からは不明。

6. 次に読むべき論文は?

要旨で参照/比較されている研究は明示されていない。関連手法として、MLLMをfine-tuningしてナビゲーションactionを予測する既存手法、およびHM3Dを用いたナビゲーション研究が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jinkai Zhang, Jingyi Xu, Yuanhong Yu, Jiarui Guo, Ruizhen Hu, Hujun Bao, Xiaowei Zhou, Sida Peng

分類: cs.RO, cs.CV

原文アブストラクト

General-purpose service robots need navigation systems that can handle diverse human requests in unfamiliar environments, combining task generality with scene generality. Some existing methods fine-tune multimodal large language models (MLLMs) to predict navigation actions, making their behavior dependent on the coverage of navigation training data and potentially limiting generalization to new requests and environments. Our key insight is to let the MLLM focus on interpreting requests, understanding scenes, and making decisions while preserving its general-purpose capabilities and delegating motion execution to navigation tools. To realize this idea, we introduce SuperNav, which equips a pretrained MLLM with a specialized agent harness without navigation-specific fine-tuning of the MLLM. Our harness supports these decisions with Navigation Skills, agent-oriented Tools for physical interaction, and task-progress and context management. A unified visual-point interface connects decision-making to motion by allowing the model to specify destinations directly in images and revise its decisions from execution feedback. Together, these components support sustained navigation across different task requirements and environments. SuperNav outperforms four evaluated baselines on instance-level, multi-object, and demand-driven tasks. Category-level evaluation on HM3D and deployment on a real quadruped robot further demonstrate its applicability across environments. Project Page: https://zju3dv.github.io/SuperNav/

関連論文

PR本紙発行元 EmplifAI