日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ベンチマークarXiv:2608.14721

AeroGround: 航空・地上協調推論のための包括的ベンチマーク

AeroGround: A Comprehensive Benchmark for Aerial-Ground Collaborative Reasoning

シェア:XThreadsFacebookLINEはてブBluesky

UAVと地上の協調シナリオにおける視覚言語モデルの推論能力を評価する新しいベンチマークを提案し、既存モデルと人間の性能差を明らかにした。

詳しい要約

1. どんなもの?

AeroGroundは、空中と地上の協調推論(aerial-ground collaborative reasoning)におけるVision-Language Models (VLMs)の評価を目的とした包括的なベンチマークである。シミュレーション環境で生成された約29,000のマルチモーダル観測グループから構成され、2,250の高品質な質問応答インスタンスを含む。質問は、クロスビュー対応、空間理解、推論の3つのカテゴリに分類される。既存のUAVベンチマークが主に空中視点に焦点を当てているのに対し、AeroGroundは現実の応用(救助、インフラ点検など)で重要となる空中と地上の協調シナリオに焦点を当てている。

2. 先行研究と比べてどこがすごい?

既存のUAVベンチマークは主に空中視点のシナリオに焦点を当てており、空中と地上の協調シナリオにおけるVLMの性能は未検証であった。AeroGroundは、このギャップを埋めるために、空中と地上の協調推論に特化した初の包括的ベンチマークを提供する点が優れている。また、16の事前学習済みVLMと2つのドメイン適応変種を評価し、人間の性能(93.3%)と最良モデル(54.4%)の間に大きなギャップがあることを明らかにしている。

3. 技術・手法の肝は?

AeroGroundは、シミュレーション環境で多様なオープン環境から約29,000のマルチモーダル観測グループを収集し、それに基づいて2,250の質問応答インスタンスを構築している。質問は、クロスビュー対応、空間理解、推論の3つのカテゴリに分類され、空中と地上の視点を組み合わせた協調推論を評価する。評価には、16の事前学習済みVLMと2つのドメイン適応変種を使用し、各モデルの精度を測定している。

4. どうやって有効だと検証した?

16の事前学習済みVLMと2つのドメイン適応変種をAeroGround上で評価した。最良モデルは平均精度54.4%を達成したが、人間の性能は93.3%であり、大きなギャップが確認された。これにより、現在のVLMが空中と地上の協調推論において限界があることを示した。

5. 議論はある?

要旨からは、現在のVLMと人間の性能の差が大きいことが示されており、空中と地上の協調推論の難しさが浮き彫りになっている。また、既存のVLMがこのタスクでどのような強みと限界を持つかを体系的に明らかにしているが、具体的な議論の内容(例えば、モデルの種類による性能差の原因など)は要旨からは不明である。

6. 次に読むべき論文は?

要旨で参照されている関連研究は明示されていないが、同分野の定番として、UAV視点のVLMベンチマークや、マルチモーダル推論のベンチマーク(例えば、一般的なVQAベンチマーク)が考えられる。具体的には、UAV向けの既存のベンチマークや、VLMの推論能力を評価する研究が関連する。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Shenghong Yi, Lin Zhang, Muzian Li, Jiakang Yuan, Haoyu Zhang, Peng Ye, Jiayuan Fan, Huafeng Qin, Tao Chen

分類: cs.CV

原文アブストラクト

Vision-language models (VLMs) have been widely employed in understanding and reasoning tasks for unmanned aerial vehicles (UAVs). Existing UAV benchmarks primarily focus on aerial-view scenarios. However, whether current VLMs can perform well on understanding and reasoning tasks in aerial-ground collaborative scenarios which are practical in real-world applications like rescue and infrastructure inspection remains underexplored. To address this gap, we introduce AeroGround, a comprehensive benchmark for evaluating VLMs in aerial-ground collaborative reasoning. AeroGround is built upon a simulated aerial-ground dataset containing approximately 29,000 multimodal observation groups from diverse open environments, and provides 2,250 high-quality question-answering instances covering cross-view correspondence, spatial understanding, and reasoning. Experiments on 16 pretrained VLMs, together with two domain-adapted variants, reveal a substantial gap between current models and human performance: the best model achieves an average accuracy of 54.4%, whereas humans reach 93.3%. By systematically revealing the strengths and limitations of existing models in aerial-ground collaborative reasoning, AeroGround provides a foundation for developing more capable aerial-ground collaborative embodied intelligence systems.

関連論文