日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.39915

NavHarness: エージェント型視覚言語ナビゲーションのための適応的ゴール設定

NavHarness: Adaptive Goals for Agentic Vision-Language Navigation

シェア:XThreadsFacebookLINEはてブBluesky

視覚言語ナビゲーションにおいて、ゴール設定・検証・記憶圧縮・視覚運動実行の4エージェントで構成されるフレームワークを提案し、長期的タスクでの一貫性と推論効率を改善した。

詳しい要約

1. どんなもの?

- 本論文は、Vision-Language Navigation (VLN) のためのAgentic VLNフレームワーク「NavHarness」を提案する。 - 4つのエージェント(Goal Agent, Verify Agent, Memory Agent, Visuomotor Agent)から構成される。 - Goal Agentは指示・観測・実行履歴に基づき適応的なゴールを設定し、Visuomotor Agentがそのゴールを達成するためのナビゲーション行動を実行する。 - Verify Agentはゴール固有の検証質問を用いて動的に完了を判定し、Memory Agentは検証済みのゴール完了を境界としてマルチモーダルな相互作用履歴を圧縮する。 - R2R-CEとRxR-CEで評価し、実世界でも8つの難ルートで成功率83.3%、ナビゲーション誤差1.51mを達成した。

2. 先行研究と比べてどこがすごい?

- 従来の汎用マルチモーダルエージェントは、局所的な行動選択がもっともらしくても、長期的なタスクでは意図した経路との整合性が保証されない問題があった。 - また、蓄積された相互作用履歴が意思決定の入力を増大させ、推論オーバーヘッドが大きくなる課題があった。 - 本手法は、適応的ゴール設定と動的検証、および履歴圧縮を組み合わせることで、これらの問題に対処している。 - 具体的な先行研究との比較は要旨からは不明。

3. 技術・手法の肝は?

- Goal Agentが指示・現在の観測・実行履歴に基づいて適応的ゴールを策定する。 - Visuomotor Agentが各ゴールを達成するためのナビゲーション行動を実行する。 - Verify Agentがゴール固有の検証質問を用いて、観測結果が意図した完了条件を満たすかを動的に評価する。 - 検証済みのゴール完了がMemory Agentによるマルチモーダル相互作用履歴の圧縮の境界となり、後続のナビゲーションに必要な情報を保持する。

4. どうやって有効だと検証した?

- R2R-CEとRxR-CEでナビゲーションを評価した。 - 3つのモデルバックボーンにわたるフレームワークの変種を検討した。 - 実行中のコンテキストの変化を調査した。 - 実世界評価では、8つの挑戦的なルートを各3回評価し、成功率83.3%、ナビゲーション誤差1.51mを達成した。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 同分野の定番として、Vision-Language Navigation (VLN) の関連手法(例:R2R-CE, RxR-CEを用いた研究)を挙げる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Haoxiang Shi, Zaijing Li, Muhe Ding, Xiang Deng, Yaowei Wang, Liqiang Nie

分類: cs.CV, cs.RO

原文アブストラクト

Vision-Language Navigation (VLN) requires embodied agents to generate actions based on instructions and observations. General-purpose multimodal agents offer a promising basis for this task, but selecting plausible local actions does not ensure that execution remains consistent with the intended route, particularly in long-horizon tasks. Moreover, the accumulated interaction history increases the input required for subsequent decisions, resulting in a significant inference overhead. To this end, we introduce \method, an Agentic VLN framework that includes a Goal Agent that sets adaptive goals for local actions, a Verify Agent that dynamically verifies whether a goal has been completed, a Memory Agent for multimodal context compression, and a Visuomotor Agent to execute adaptive goals. Specifically, the Goal Agent formulates adaptive goals based on the instruction, current observation, and execution history. Then the Visuomotor Agent executes navigation actions to achieve each goal, while the Verify Agent uses a goal-specific verification question to dynamically assess whether the observed outcomes satisfy the intended completion condition. Verified goal completion then marks a boundary for the Memory Agent to compress the corresponding multimodal interaction history while preserving information needed for subsequent navigation. We evaluate navigation on R2R-CE and RxR-CE, examine framework variants across three model backbones, and study context evolution during execution. For Real-World evaluation, \method achieves 83.3\% success and 1.51\,m navigation error across eight challenging routes evaluated three times each.

関連論文

PR本紙発行元 EmplifAI