ドローン画像におけるゼロショット人物検出と行動認識のためのYOLO-WorldとGPT-4V LMMの活用
Leveraging YOLO-World and GPT-4V LMMs for Zero-Shot Person Detection and Action Recognition in Drone Imagery
ドローン視点の画像で、ゼロショット大規模マルチモーダルモデル(YOLO-WorldとGPT-4V)を用いて人物検出と行動認識を評価した。YOLO-Worldは良好な検出性能を示し、GPT-4Vは行動分類は苦手だが不要領域のフィルタリングや情景記述に有望な結果を出した。
著者: Christian Limberg, Artur Gonçalves, Bastien Rigault, Helmut Prendinger
分類: cs.CV, cs.RO
原文アブストラクト
In this article, we explore the potential of zero-shot Large Multimodal Models (LMMs) in the domain of drone perception. We focus on person detection and action recognition tasks and evaluate two prominent LMMs, namely YOLO-World and GPT-4V(ision) using a publicly available dataset captured from aerial views. Traditional deep learning approaches rely heavily on large and high-quality training datasets. However, in certain robotic settings, acquiring such datasets can be resource-intensive or impractical within a reasonable timeframe. The flexibility of prompt-based Large Multimodal Models (LMMs) and their exceptional generalization capabilities have the potential to revolutionize robotics applications in these scenarios. Our findings suggest that YOLO-World demonstrates good detection performance. GPT-4V struggles with accurately classifying action classes but delivers promising results in filtering out unwanted region proposals and in providing a general description of the scenery. This research represents an initial step in leveraging LMMs for drone perception and establishes a foundation for future investigations in this area.