日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
クロスビュー検出arXiv:2608.27997v1

A-PAIR: 空対地クロスビュー人物参照検出のためのベンチマークとID一貫性グラウンディングフレームワーク

A-PAIR: A Benchmark and Identity-Consistent Grounding Framework for Air-Ground Cross-View Referring Person Detection

シェア:XThreadsFacebookLINEはてブBluesky

空対地のクロスビューで人物を言語で参照して検出する新しいタスクを定義し、ベンチマークとID一貫性を保つグラウンディング手法を提案した。

詳しい要約

1. どんなもの?

A-PAIRは、Air-Ground Cross-View Referring Person Detection (AGCV-RPD) のための最初の包括的ベンチマークであり、22,137のクロスビュー参照サンプルを含む。また、Identity-Consistent Referring Grounding (ICRG) というフレームワークを提案し、因子化された参照グラウンディング、候補完全性監視、クロスビュー一貫性キャリブレーションを組み合わせて、空中と地上のエージェント間で同一人物を特定する。

2. 先行研究と比べてどこがすごい?

既存のReferring Expression ComprehensionやOpen-Vocabulary Grounding手法は、クロスビューのID一貫性を考慮しておらず、類似する歩行者妨害、弱い空中外観手がかり、クロスビューID一貫性を必要とするAGCV-RPDには不十分である。A-PAIRはこの問題に特化した最初のベンチマークであり、ICRGはペアレベルでの検出を改善する。

3. 技術・手法の肝は?

ICRGは、因子化された参照グラウンディング(factorized referential grounding)、候補完全性監視(candidate-completeness supervision)、クロスビュー一貫性キャリブレーション(cross-view consistency calibration)を組み合わせて、空中と地上のペア選択を共同で行う。また、A-PAIRの構築には、Factorized Annotation and Referential Alignment (FARA) という半自動アノテーションフレームワークを用い、コストを削減しつつ因子化された参照記述とID一貫性監視を生成する。

4. どうやって有効だと検証した?

ICRGを強力なベースラインと比較し、地上、空中、ペアレベルの検出で改善を示した。特に、ペアF1を16.65%から22.28%に向上させた。

5. 議論はある?

要旨からは、AGCV-RPDにはペア検出とID一貫性推論が必要であることが示唆されるが、具体的な議論や限界については不明。

6. 次に読むべき論文は?

要旨で参照されている既存のReferring Expression ComprehensionやOpen-Vocabulary Grounding手法、および関連するクロスビュー人物検出の研究。具体的には、Referring Expression ComprehensionのベンチマークやOpen-Vocabulary Groundingの手法が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhoupeng Guo, Xinjie Yao, Yunqi Zhu, Zhihe Fan, Siqi Zhao, Jianjun Chen, Yichen Dong, Yan Fan, Pengfei Zhu

分類: cs.CV, cs.MM

原文アブストラクト

Air-ground cross-view referring person detection is a necessary component in the language-to-perception-to-control chain of collective embodied intelligence, grounding a language command into the same physical target before ground and aerial agents can coordinate downstream actions. Existing referring expression comprehension and open-vocabulary grounding methods do not jointly account for cross-view identity consistency, making them insufficient for Air-Ground Cross-View Referring Person Detection (AGCV-RPD), which involves similar pedestrian distractors, weak aerial appearance cues, and cross-view identity consistency. To study this problem, we introduce Air-Ground Paired Identity-Aware Referring (A-PAIR), the first comprehensive AGCV-RPD benchmark, containing 22,137 cross-view referring samples. To construct A-PAIR efficiently, we propose Factorized Annotation and Referential Alignment (FARA), a semi-automatic annotation framework that generates factorized referring descriptions and identity-consistency supervision at reduced cost. We propose Identity-Consistent Referring Grounding (ICRG), a framework that combines factorized referential grounding, candidate-completeness supervision, and cross-view consistency calibration for joint air-ground pair selection. ICRG improves ground, aerial, and pair-level detection over strong baselines, increasing pair F1 from 16.65% to 22.28%. These results show that AGCV-RPD requires paired detection and identity-consistent reasoning.