日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.09764

AeroEval: AI生成ドローン任務のための段階的プログラム・実行検証

AeroEval: Staged Program and Execution Validation for AI-Generated Drone Missions

シェア:XThreadsFacebookLINEはてブBluesky

LLMが生成したドローン任務プログラムを、構文・API・意図・実行軌跡の段階で検証し、失敗箇所を特定して再生成を促すミドルウェア。ナビゲーション成功率を55%から95%に改善。

詳しい要約

1. どんなもの?

- AI生成ドローンミッションの段階的検証ミドルウェアAeroEvalを提案 - LLMが自然言語から生成したドローンコードは文法的に正しくても意図・環境制約・ミッション挙動に違反しうる - 決定論的プログラム解析と文脈接地型LLMエージェントを組み合わせる - 構文・API使用・ミッション意図を検証後、実行軌跡・要件・環境文脈で挙動を評価 - 各段階が構造化失敗情報を返し反復再生成を可能にする

2. 先行研究と比べてどこがすごい?

- 既存のドローンコード生成系はprompt guardrailsやシミュレータ結果に依存し失敗局所化が限定的 - AeroEvalはprogram-levelとexecution-groundedのagentic検証を統合 - 段階的検証により失敗の局所化と反復再生成を実現 - ナビゲーション成功率を55%から95%へ改善 - 解析ミッションでone-shot AeroGenの34%から88%へ集約成功率を向上

3. 技術・手法の肝は?

- 決定論的プログラム解析とcontext-grounded LLM agentsを組み合わせたagent-assisted middleware - 第1段階でprogram syntax、platform API usage、mission intentを検証 - 第2段階でexecution trajectories、mission requirements、environmental contextを用いて実現挙動を評価 - 各段階がstructured failure informationを返しiterative regenerationを駆動 - Code ValidatorとTrajectory Validatorを含むパイプライン構成

4. どうやって有効だと検証した?

- AirSimとGazeboシミュレータ上で20 navigation tasksと5 analytical mission typesを評価 - ナビゲーション成功率が55%から95%へ向上 - stagewise ablationでCode Validator単体44%、Trajectory Validator単体56%、フルパイプライン88% - 段階がprogram structure、API usage、mission intent、obstacle avoidance、altitude、coverage、event-driven transitionsの相補的失敗を検出 - 解析ミッションでone-shot AeroGenの34%から88%へ改善

5. 議論はある?

- 評価環境におけるprogram-levelとexecution-grounded agentic validationの組み合わせの利点を実証 - 段階的検証が相補的失敗を検出しguided regenerationが修正 - 失敗局所化の限界を改善 - ただし評価はAirSimとGazeboシミュレータに限定 - 実世界展開や他環境への一般化は要旨からは不明

6. 次に読むべき論文は?

- AeroGen(one-shotベースラインとして比較) - AirSimおよびGazeboを用いたドローンコード生成研究 - LLM-based code generationの検証手法 - prompt guardrailsやsimulator outcomesに基づく既存ドローンコード生成システム

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Kautuk Astu, Naina Rabha, Yogesh Simmhan

分類: cs.RO, cs.DC

原文アブストラクト

Large Language Models (LLMs) can generate drone programs from natural-language mission descriptions, but syntactically valid programs may still violate user intent, environmental constraints, and mission-level behavior. This problem is pronounced in cyber-physical applications, where correctness depends on the interaction among generated code, mobile sensing, environmental geometry, event-driven analytics, and physical execution. Existing drone code-generation systems primarily use prompt guardrails or simulator outcomes and provide limited failure localization. We present AeroEval, an agent-assisted middleware for staged validation of AI-generated drone missions. AeroEval combines deterministic program analysis with context-grounded LLM agents. It first validates program syntax, platform API usage, and mission intent, and then evaluates the realized behavior using execution trajectories, mission requirements, and environmental context. Each stage returns structured failure information for iterative regeneration. In our evaluation using 20 navigation tasks and five analytical mission types over AirSim and Gazebo simulators, AeroEval improves navigation success from 55% to 95%. In a stagewise ablation study, our Code and Trajectory Validators by themselves achieve mean run-level success rates of 44% and 56%, respectively, while the full AeroEval pipeline achieves 88%; the stages detect complementary failures in program structure, API usage, mission intent, obstacle avoidance, altitude, coverage, and event-driven transitions and the guided regeneration corrects for them. Across the main analytics missions, AeroEval increases aggregate run-level success from 34% for one-shot AeroGen to 88% within the regeneration budget. These results demonstrate the benefit of combining program-level and execution-grounded agentic validation for AI-generated drone applications in the evaluated environment.

関連論文

PR本紙発行元 EmplifAI