日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
航空探索arXiv:2609.36066

AerialDojo-200K:オープンワールド航空物体目標探索のための大規模ベンチマークスイート

AerialDojo-200K: A Large-Scale Benchmark Suite for Open-World Aerial Object-Goal Search

シェア:XThreadsFacebookLINEはてブBluesky

航空エージェントが大規模3D環境を自律探索し、意味記述や参照画像で指定された物体に到達するタスクのための、既存比3倍のシーンと18.7倍のタスクを含む大規模ベンチマークを構築した。

詳しい要約

1. どんなもの?

- 開放世界の航空物体目標探索のための大規模ベンチマークスイート。 - 意味記述や参照画像で指定された目標物体に到達するタスク。 - 42のシミュレーションシーン(4つのシーンファミリー、21のシーンタイプ)を含む。 - 205,732のタスクインスタンス(Base, Standard, Long-Horizon設定)。 - 各インスタンスに衝突のない参照軌道とマルチビュー動画記録が付属。

2. 先行研究と比べてどこがすごい?

- 既存の最大ベンチマークと比較して、シーン数が3倍、タスクインスタンス数が18.7倍。 - 従来は小規模で環境固有のベンチマークが多く、行動空間やデータ形式が異質だった。 - 大規模な訓練とクロスベンチマーク評価を可能にし、スケーラビリティと一般化性を向上。

3. 技術・手法の肝は?

- 42のシミュレーションシーンを構築(都市18、自然12、インフラ6、災害6)。 - 12名のアノテータが2ヶ月かけて手動で109のランドマーク、2099の目標物体、2099の物体アンカーを注釈。 - 205,732のタスクインスタンスを構築(意味目標100K以上、画像目標100K以上)。 - 各タスクに衝突のない参照軌道とマルチビュー動画を提供。 - 統一評価フレームワークを開発(21のin-distributionシーンと21のout-of-distributionシーン)。

4. どうやって有効だと検証した?

- 5つのオープンソースと4つのクローズドソースのマルチモーダル大規模言語モデルを評価。 - 評価の結果、汎用航空エージェントの実現にはまだ長い道のりがあることが明らかになった。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 同分野の定番として、AerialVLN、Open-World Object-Goal Navigation、Multimodal Large Language Models for Navigationなどが考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Tongtong Feng, Xin Wang, Haoran Hou, Ren Wang, Weiran Wang, Shaokai Zhu, Ziqi Jia, Hao Wang, Yu-Wei Zhan, Zongyuan Wu, Jinghao Cui, Wenwu Zhu

分類: cs.CV, cs.AI, cs.ET, cs.MM, cs.RO

原文アブストラクト

Open-world aerial object-goal search is a foundational yet challenging task, requiring aerial agents to autonomously explore large-scale, unstructured three-dimensional environments and reach target objects specified by semantic descriptions or reference images, rather than following route-specific instructions. However, research in this task remains at a nascent stage and relies on small, environment-specific benchmarks with heterogeneous action spaces and data formats. These limitations hinder large-scale training and cross-benchmark evaluation, constraining the scalability and generalizability of aerial agents. To address this problem, we propose AerialDojo-200K, a large-scale benchmark suite for open-world aerial object-goal search, with 3 times as many scenes and 18.7 times as many task instances as the largest existing benchmark for this task. Specifically, we construct 42 simulation scenes spanning four scene families and 21 scene types, including 18 urban, 12 natural, six infrastructure, and six disaster scenes. To ensure data quality, 12 annotators spent two months manually annotating 109 landmarks, 2099 target objects, and 2099 object anchors across these scenes. We further construct 205,732 task instances, comprising over 100K semantic-goal and over 100K image-goal instances across Base, Standard, and Long-Horizon settings. Each task instance includes a collision-free reference trajectory and corresponding multi-view video recordings. We also develop a unified evaluation framework with a scene partition comprising 21 in-distribution scenes and 21 out-of-distribution scenes. Finally, our evaluation of five open-source and four closed-source multimodal large language models reveals that there is still a long way to go toward achieving general-purpose aerial agents. All can be found at https://fengtt42.github.io/AerialDojo/.

PR本紙発行元 EmplifAI