日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ナビゲーションarXiv:2609.20116

GPT-6-Astraによる連続環境でのゼロショット視覚言語ナビゲーションの行動分析

GPT-6-Astra in a Navigation Workflow: Behavioral Analysis in Zero-Shot Vision-and-Language Navigation in Continuous Environments

シェア:XThreadsFacebookLINEはてブBluesky

GPT-6-Astraをゼロショットの視覚言語ナビゲーションに適用し、指示理解と行動提案の挙動を分析した。成功率52%だが、局所判断と自律完遂の間にギャップがあることを示した。

詳しい要約

1. どんなもの?

GPT-6-Astraをzero-shot VLN-CEシステムに適用し、その挙動を分析した研究。 - システムはobservation–decision–executionワークフローを採用。 - パッケージ化されたagent harnessやナビゲーション特化のfine-tuningは使用しない。 - 各リクエストは選択された観測、実行フィードバック、保持された進捗記録を受け取る。 - 評価はcontext managementとaction controlを含むシステム全体を対象。

2. 先行研究と比べてどこがすごい?

Open-Navが用いたR2R-CE val-unseenの100エピソードのうち50で評価。 - 成功率52.0%、SPL 48.9%、nDTW 70.8%を達成。 - 先行研究との直接比較や優位性の主張は要旨からは不明。 - zero-shotかつfine-tuningなしで動作する点が特徴。

3. 技術・手法の肝は?

observation–decision–executionワークフローと直接的なmodel API呼び出し。 - 各リクエストに選択された観測、実行フィードバック、保持された進捗記録を入力。 - context managementとaction controlを含むシステム全体を評価。 - ナビゲーション特化のfine-tuningやagent harnessは不使用。

4. どうやって有効だと検証した?

R2R-CE val-unseenの100エピソード中50で評価。 - 成功率52.0%、SPL 48.9%、nDTW 70.8%を報告。 - 記録された応答の分析により3つの知見を提示。 - 観測と提供された履歴を用いてランドマークと以前の行動を指示に結びつける。 - 追加ビューの要求や不確実な判断の修正を含むレビュー。 - タスク理解と自律的完了の間のギャップを示唆。 - 終了時、36.0%のエピソードがworkflow-accepted STOPで成功、別の16.0%がステップ制限時に距離基準を満たす。

5. 議論はある?

中心的な課題として、正しい局所判断を持続的な進捗と適切な停止に変換することを指摘。 - 未完了の横断が認識されつつ回転が続く例を報告。 - タスク理解と自律的完了の間にギャップがあると示唆。 - その他の議論や限界は要旨からは不明。

6. 次に読むべき論文は?

Open-Nav(R2R-CE val-unseenの100エピソードを使用した先行研究)。 - 関連手法としてVLN-CE、zero-shot Vision-and-Language Navigation、R2R-CEデータセット。 - その他の具体的な論文は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Guangzhao Dai, Qi Wu, Bin Zhu

分類: cs.RO

原文アブストラクト

We study GPT-6-Astra in a zero-shot Vision-and-Language Navigation in Continuous Environments (VLN-CE) system, where it interprets instructions, assesses its surroundings, and proposes actions. The system uses a common observation--decision--execution workflow with direct model API calls, without a packaged agent harness or navigation-specific fine-tuning. In this workflow, each request receives selected observations, execution feedback, and retained progress records. Evaluation covers the complete system, including context management and action control. We evaluate the system on 50 of the 100 R2R-CE val-unseen episodes used by Open-Nav. It achieves a success rate of 52.0\%, an SPL of 48.9\%, and an nDTW of 70.8\%. Our analysis highlights three findings. First, recorded responses link landmarks and earlier actions to instructions using observations and supplied history. Second, reviews include requests for additional views and revisions of uncertain judgments. Third, the results suggest a gap between task understanding and autonomous completion: an unfinished crossing is recognized while rotation continues. At termination, 36.0\% of episodes succeed with a workflow-accepted STOP, while another 16.0\% meet the distance criterion at the step limit. These results highlight a central challenge: translating correct local judgments into sustained progress and appropriate stopping.

関連論文

PR本紙発行元 EmplifAI