日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2405.16994

視覚と言語によるナビゲーションのための生成事前学習トランスフォーマー

Vision-and-Language Navigation Generative Pretrained Transformer

シェア:XThreadsFacebookLINEはてブBluesky

GPT2デコーダで軌跡系列の依存関係をモデル化し、履歴エンコーダを不要にしたVLN手法を提案。模倣学習による事前学習と強化学習による微調整を分離し、既存の複雑なエンコーダベース手法を上回る性能を達成した。

著者: Wen Hanlin

分類: cs.AI, cs.CL, cs.CV, cs.RO

原文アブストラクト

In the Vision-and-Language Navigation (VLN) field, agents are tasked with navigating real-world scenes guided by linguistic instructions. Enabling the agent to adhere to instructions throughout the process of navigation represents a significant challenge within the domain of VLN. To address this challenge, common approaches often rely on encoders to explicitly record past locations and actions, increasing model complexity and resource consumption. Our proposal, the Vision-and-Language Navigation Generative Pretrained Transformer (VLN-GPT), adopts a transformer decoder model (GPT2) to model trajectory sequence dependencies, bypassing the need for historical encoding modules. This method allows for direct historical information access through trajectory sequence, enhancing efficiency. Furthermore, our model separates the training process into offline pre-training with imitation learning and online fine-tuning with reinforcement learning. This distinction allows for more focused training objectives and improved performance. Performance assessments on the VLN dataset reveal that VLN-GPT surpasses complex state-of-the-art encoder-based models.

関連論文

PR本紙発行元 EmplifAI