日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.08526

WareFly-VLA: スマート倉庫におけるUAVナビゲーションと人物追跡のための視覚-言語-行動フレームワーク

WareFly-VLA: A Vision-Language-Action Framework for UAV Navigation and Human Tracking in Smart Warehouses

シェア:XThreadsFacebookLINEはてブBluesky

倉庫内で言語指示に従いUAVが人物を探索・追跡するためのフォトリアルなVLAデータセットとベンチマークを構築し、4つの既存VLAモデルを評価した。

詳しい要約

1. どんなもの?

- スマート倉庫での UAV ナビゲーションと人間追跡のための VLA フレームワークおよびデータセット - 言語条件付きの人間探索・位置推定・追跡を対象 - NVIDIA Isaac Sim で収集した 507 の人間テレオペ飛行エピソードと 8,504 の高解像度 RGB 遷移を含む - 各遷移に人間が書いた外観記述と同期した 4 自由度制御コマンドを付与 - タスクは target approach と person following の 2 つ - 遮蔽、長距離探索、高度変化、 clutter を含む - 4 つのオープンソース VLA アーキテクチャの統一ベンチマークを提供 - データセット、ベースライン、評価プロトコルを公開

2. 先行研究と比べてどこがすごい?

- 従来の VLA はロボットマニピュレーションや地上移動ナビゲーションで成果 - スマート倉庫における UAV の言語条件付き制御はほとんど未探索 - 連続低レベル飛行行動、細粒度自然言語目標記述、現実的産業環境を同時に提供するベンチマークが欠如 - 本研究はこれらを統合した初のフォトリアリスティック UAV VLA フレームワークとデータセットを提供 - 漏洩のないエピソードレベルプロトコルで 2 つの制御レート下に統一ベンチマークを確立 - 地上・ヒューマノイドから航空プラットフォームへの foundation-model インターフェースの転移が不十分であることを示す

3. 技術・手法の肝は?

- NVIDIA Isaac Sim を用いたフォトリアリスティックな倉庫環境 - 人間テレオペレーションによる飛行エピソード収集 - 各遷移に同期した 4 自由度制御コマンドを付与 - 人間が書いたターゲット作業者の外観記述を言語条件として使用 - 4 つの VLA アーキテクチャ(SmolVLA, GR00T N1.7, pi_0, OpenVLA)を統一評価 - 漏洩のないエピソードレベルプロトコルを採用 - 2 つの制御レートで評価 - 連続行動モデリングと離散行動トークン化を比較 - 同期された video, language, action, pose, difficulty アノテーションを提供

4. どうやって有効だと検証した?

- 4 つのオープンソース VLA アーキテクチャを統一ベンチマークで評価 - 漏洩のないエピソードレベルプロトコルと 2 つの制御レートで検証 - 厳密な汎化設定下で性能が大幅に低下することを確認 - 連続行動モデリングが離散行動トークン化より一貫して優れることを示す - 単一フレームから信頼性高く学習できるのは forward チャネルのみであることを確認 - 現在の foundation-model インターフェースが地上・ヒューマノイドから航空プラットフォームへ転移しにくいことを示す

5. 議論はある?

- 倉庫における言語条件付き航空制御は未解決であると結論 - 厳密な汎化設定下での性能低下が課題 - 連続行動モデリングの優位性が示唆される - 単一フレームからの学習可能性は forward チャネルに限定 - foundation-model インターフェースの embodiment 間転移が不十分 - 同期された video, language, action, pose, difficulty アノテーションが world-model 研究を支援 - データセット、ベースライン、評価プロトコルを公開し、言語に基づく航空自律性を支援

6. 次に読むべき論文は?

- SmolVLA - GR00T N1.7 - pi_0 - OpenVLA - 要旨で参照/比較されている研究や関連手法を挙げる(上記 4 つ)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Thinh D. Le, Son T. Nguyen, Duong Q. Nguyen, Dung D. Le, Ngo Anh Vien, H. Nguyen-Xuan

分類: cs.RO, cs.CV

原文アブストラクト

Vision-Language-Action (VLA) models have achieved impressive results in robotic manipulation and ground-mobile navigation, yet language-conditioned control of unmanned aerial vehicles (UAVs) in smart warehouses remains largely unexplored, hindered by the lack of benchmarks that jointly provide continuous low-level flight actions, fine-grained natural-language target descriptions, and realistic industrial environments. This paper introduces WareFly-VLA, a photorealistic UAV VLA framework and dataset for language-guided human search, localization, and tracking in warehouse environments. It contains 507 human-teleoperated flight episodes and 8,504 high-resolution RGB transitions collected in NVIDIA Isaac Sim, each paired with a human-written appearance description of the target worker and a synchronized four-degree-of-freedom control command. Two aerial tasks are covered: target approach and person following, under occlusion, long-range search, altitude variation, and clutter. A unified benchmark of four open-source VLA architectures (SmolVLA, GR00T N1.7, pi_0 and OpenVLA) is established under a leakage-free episode-level protocol at two control rates. The results show that language-conditioned aerial control in warehouses is far from solved: performance drops substantially under strict generalization settings, continuous action modeling consistently outperforms discrete action tokenization, only the forward channel is reliably learnable from a single frame, and current foundation-model interfaces transfer poorly from ground and humanoid embodiments to aerial platforms. The synchronized video, language, action, pose, and difficulty annotations further support world-model research. The dataset, baselines, and evaluation protocol are released to support language-grounded aerial autonomy in smart warehouses.

関連論文

PR本紙発行元 EmplifAI