日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.37922

WayFinder: ゼロショット経由点生成と低レベル運動制御のための階層的視覚-言語-行動フレームワーク

WayFinder: Hierarchical Visual-Language-Action for Zero-Shot Waypoint Generation and Low-Level Kinematic Control

シェア:XThreadsFacebookLINEはてブBluesky

大規模マルチモーダルモデルで高レベルな経由点をゼロショット生成し、軽量なオンボード制御と非同期に組み合わせることで、ファインチューニング不要でロボットナビゲーションの成功率を最大27.45%向上させた。

詳しい要約

1. どんなもの?

- Visual Language Action (VLA) モデルの実世界展開を阻む微調整コストと実行信頼性の問題を解決する - 微調整不要の end-to-end 閉ループ階層型 VLA フレームワーク WayFinder を提案 - 高レベルタスク推論と低レベル運動制御を分離 - ゼロショットの offboard Multimodal Large Language Model (MLLM) が言語文脈と状態マップから戦略的 waypoint を生成 - 軽量な onboard ポリシーが連続センサフィードバックに基づき高頻度で実時間運動制御を非同期実行

2. 先行研究と比べてどこがすごい?

- 従来の VLA は特定ロボット実装やタスクへの微調整が必須で計算コストが高い - WayFinder は微調整を不要にし、高レベル MLLM をナビゲーション失敗時のみ呼び出す - 高価な推論を最小化しつつ、ナビゲーション成功率を最大 27.45% 向上 - ベースラインの低レベルポリシーと比較して優れたナビゲーション信頼性を実証

3. 技術・手法の肝は?

- 階層型アーキテクチャで高レベル推論と低レベル制御を分離 - 高レベル: ゼロショット offboard MLLM ポリシーが言語文脈と状態マップを処理し waypoint を生成 - 低レベル: 軽量 onboard ポリシーが連続センサフィードバックに基づき高頻度で実時間運動制御 - 非同期実行により計算負荷を分散 - ナビゲーション失敗時のみ高レベル MLLM をクエリし、微調整を排除

4. どうやって有効だと検証した?

- Microsoft AirSim で評価 - 複雑度の異なる 4 環境でテスト - 3 つの MLLM スケールを比較し、予測性能と計算効率のバランスを検証 - ベースライン低レベルポリシーと比較してナビゲーション信頼性が優れることを確認 - 成功率が最大 27.45% 向上

5. 議論はある?

- 微調整不要で計算コストを削減しつつ成功率を向上 - 高レベル MLLM の呼び出しを失敗時に限定することで効率化 - ただし要旨からは、失敗検出の閾値や一般化限界、実機展開の課題についての議論は不明 - シミュレーション評価のみで実世界実験の有無は要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない - 関連手法として Visual Language Action (VLA) モデル、Multimodal Large Language Model (MLLM) を用いたナビゲーション、階層型ロボット制御、ゼロショット waypoint 生成の定番研究を挙げる - 具体的な論文名は要旨からは不明

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Timothy K Johnsen, Marco Levorato

分類: cs.RO

原文アブストラクト

Visual Language Action (VLA) models offer unprecedented generalization for autonomous robots; however, their real-world deployment is frequently bottlenecked by unreliable execution and the prohibitive computational cost of fine-tuning for specific robot embodiments and tasks. To bridge this gap, we propose WayFinder, an end-to-end, closed-loop hierarchical VLA framework that circumvents the need for fine-tuning by decoupling high-level task reasoning from low-level kinematic control. WayFinder utilizes a zero-shot, offboard Multimodal Large Language Model (MLLM) policy to process linguistic context and state maps for strategic waypoint generation. Asynchronously, a lightweight, onboard policy executes real-time kinematic control at high frequency based on continuous sensor feedback. We evaluate WayFinder in Microsoft AirSim, testing on four environments of varying complexity and three MLLM scales to balance prediction efficacy with computational efficiency. Our results demonstrate that WayFinder achieves superior navigation reliability compared to baseline low-level policies. By querying the high-level MLLM only during navigation failures, WayFinder eliminates the need for fine-tuning, minimizes expensive inferences, and significantly increases navigation success rates by up to 27.45%.

関連論文

PR本紙発行元 EmplifAI