日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.05024

LightVLN:コンパクトメモリと履歴誘導型局所集約による効率的な空中視覚言語ナビゲーション

LightVLN: Efficient Aerial Vision-and-Language Navigation with Compact Memory and History-Guided Local Aggregation

シェア:XThreadsFacebookLINEはてブBluesky

0.5Bの小型言語モデルと履歴・現在観測の圧縮表現を組み合わせ、ドローン上でのリアルタイム動作を可能にした軽量な空中VLNフレームワークを提案。

詳しい要約

1. どんなもの?

- 軽量な空中Vision-and-Language Navigation (VLN) フレームワーク - 0.5Bの言語backboneと圧縮された履歴・現在観測表現を組み合わせる - 最大16履歴フレームで観測由来トークンは最大48 - OpenFlyやAerialVLN-Sで評価、DJI M350 RTK上でHIL評価

2. 先行研究と比べてどこがすごい?

- 従来の空中VLNは大規模vision-language backboneと密な視覚履歴に依存 - 計算・メモリコストが大きくオンボード展開が困難 - LightVLNは0.5B backboneで7B backboneベースラインを多くの指標で上回る - 軽量かつ高性能を両立

3. 技術・手法の肝は?

- 各履歴フレームをpolicyが既に計算した視覚特徴で単一トークンに圧縮 - history- and instruction-conditioned local aggregationを導入 - 現在観測を256から32視覚トークンに削減し空間情報を保持 - 最大16履歴フレームで観測由来トークンは最大48

4. どうやって有効だと検証した?

- OpenFlyデータセットでTest-Seen 50.93%、Test-Unseen 36.14% SR - AerialVLN-S Val-Seenで25.83% SR - 再構成した未見キャンパスでDJI M350 RTKとJetson Orin NX 16 GBに展開 - closed-loop onboard-compute real-to-sim HIL評価で14.61 Hz推論、11.13 Hz意思決定更新

5. 議論はある?

- 有効性と効率性を示すが、議論や限界は要旨からは不明 - 一般化や実環境での性能は要旨からは不明

6. 次に読むべき論文は?

- OpenFly - AerialVLN-S - 7B language-backbone baselines - 同分野の定番としてVision-and-Language Navigation (VLN) 関連研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yiming Zhao, Tianshun Li, Jingle He, Ruonan Chai, Xinhu Zheng

分類: cs.CV, cs.RO

原文アブストラクト

Aerial vision-and-language navigation (VLN) enables unmanned aerial vehicles to execute long-horizon natural-language instructions from visual observations in complex three-dimensional environments. However, recent aerial VLN models often rely on large-scale vision-language backbones and dense visual histories, imposing substantial computation and memory costs that hinder onboard deployment. We propose LightVLN, a lightweight history-aware aerial VLN framework that combines a compact 0.5B language backbone with compact representations of both historical and current observations. LightVLN compresses each historical frame into a single token using visual features already computed by the policy. It further introduces history- and instruction-conditioned local aggregation to reduce the current observation from 256 to 32 visual tokens while preserving navigation-relevant spatial information. With up to 16 historical frames, the policy uses at most 48 observation-derived tokens. On the public OpenFly dataset, LightVLN achieves 50.93% Test-Seen and 36.14% Test-Unseen success rates (SR), outperforming the evaluated 7B language-backbone baselines on most reported metrics. It also achieves 25.83% SR on AerialVLN-S Val-Seen. In a reconstructed unseen campus, we deploy LightVLN on a DJI M350 RTK with an external Jetson Orin NX 16 GB for closed-loop onboard-compute real-to-sim hardware-in-the-loop (HIL) evaluation, achieving 14.61 Hz model inference and 11.13 Hz end-to-end decision updates. These results demonstrate the effectiveness and efficiency of LightVLN for aerial navigation.

関連論文

PR本紙発行元 EmplifAI