日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.18451

VLM-MPPI: 自然言語を多様な飛行行動に接地する空中ナビゲーション

VLM-MPPI: Grounding Natural Language in Behaviorally Diverse Trajectories for Aerial Navigation

シェア:XThreadsFacebookLINEはてブBluesky

自然言語の指示を6つの行動条件付きMPPIプランナで生成した多様な軌道候補に結びつけ、VLMがFPV画像から候補を選ぶ階層型UAVナビゲーションを実機とシミュレーションで実証した。

詳しい要約

1. どんなもの?

- 自然言語意図と動的に実行可能な飛行行動を整合させる階層型UAVナビゲーション - 6つのbehavior-conditioned MPPI plannerを並列アンサンブル - 3D候補をFPV RGBに投影し、VLMが言語プロンプトから候補を選択 - MPPIは20Hzで再計画、PIDで追従 - NVIDIA Isaac Simと実機quadrotorで実装

2. 先行研究と比べてどこがすごい?

- 従来の言語誘導ナビは抽象意味と低レベル制御のギャップが課題 - 単なる確率的変動でなく、意図的に多様な行動モードを生成 - 言語接地を視覚的行動選択問題に変換 - VLM遅延下でも頑健な言語整合を実現 - 評価シナリオで100%タスク成功率

3. 技術・手法の肝は?

- 6つのbehavior-conditioned MPPI plannerを並列化 - mode-specific guiding costsとsampling biasesで軌道モードを誘導 - 各モードが独自の行動平均に収束 - 3D候補をonboard FPV RGBに投影 - 事前学習VLMがオーバーレイ画像と言語プロンプトから候補インデックスを非同期選択 - MPPI 20Hz再計画、PID低レベル制御

4. どうやって有効だと検証した?

- NVIDIA Isaac Simと実機quadrotor(LiDARとRGB搭載)で実験 - シミュレーションと実飛行の両方で検証 - 意味的に有意な行動多様性を確認 - VLM遅延下での頑健な言語整合 - 全モードで安全かつ再現可能な飛行 - 評価シナリオで100%タスク成功

5. 議論はある?

- VLM遅延下での言語整合の頑健性を議論 - 行動多様性の意味的有意性を議論 - 安全で再現可能な飛行を議論 - その他の議論は要旨からは不明

6. 次に読むべき論文は?

- Model Predictive Path Integral (MPPI) - Vision-Language Model (VLM) - PID制御 - NVIDIA Isaac Sim - 関連する言語誘導ナビゲーション研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Hanbing Zhang, Fangguo Zhao, Zerui Li, Xin Guan, Peng Cheng, Shuo Li

分類: cs.RO

原文アブストラクト

We present a hierarchical UAV navigation framework that aligns natural-language intent with dynamically feasible flight behaviors in cluttered indoor environments. To bridge the gap between abstract semantics and low-level control, we employ a parallelized ensemble of six behavior-conditioned Model Predictive Path Integral (MPPI) planners. Crucially, by designing mode-specific guiding costs and sampling biases, we induce distinct trajectory modes that converge to unique behavioral means, yielding a compact set of intentionally diverse candidates rather than mere stochastic variations. We project these 3D candidates onto the onboard first-person-view RGB stream, turning language grounding into a visual action selection problem. A pretrained vision--language model (VLM) asynchronously selects the candidate index given the overlaid FPV image and a natural-language prompt, while MPPI replans at 20Hz and a PID-based low-level controller tracks the selected trajectory. We implement the full pipeline in NVIDIA Isaac Sim and on a real-world quadrotor platform equipped with LiDAR and RGB sensing. Experiments in both simulation and real-world flights show semantically meaningful behavior diversity, robust language alignment despite VLM latency, and safe, repeatable flight across all modes, achieving 100% task success in our evaluated scenarios.

関連論文

PR本紙発行元 EmplifAI