一人称視点での手検出のためのマルチモーダルRGB・イベントデータセット
A Multimodal RGB and Events Dataset for Hand Detection in First-Person View
イベントカメラとRGBカメラを組み合わせた手検出用の合成データセットを構築し、既存のマルチモーダル検出アルゴリズムで性能を実証した。
著者: Bharghav Kota, Yulia Sandamirskaya
分類: cs.CV
原文アブストラクト
Existing hand detection algorithms work on images and the detection rate is restricted by the frame rate of the camera. In hand detection applications for moving robotic systems, conventional cameras cause motion blur, especially in darker lighting conditions. We can leverage the use of event-based cameras which possess a high dynamic range, high temporal resolution, and low power consumption. Recent work has shown that using a stereo setup of an event-based and a frame-based camera improves detection accuracy and the bandwidth-latency tradeoff. The main bottleneck in using event-based cameras in object detection and recognition tasks is a relatively low amount of training data. In this work, we propose a methodology and an exemplary synthetic event-based hand dataset from an egocentric, first-person view perspective. The data is synthesized from the existing RGB Egohands dataset with the v2e toolbox. Parameters of the v2e toolbox are varied to provide versions of the dataset with different lighting conditions and scales. Ground truth detections are generated with a fine-tuned YOLOv8 model which is applied to the RGB images in the Egohands dataset and interpolated on the high-temporal resolution events. We use the multi-modal dataset to perform hand detection with existing object detection algorithms which use a multi-modal setup of event and RGB cameras and demonstrate performance comparable to the state-of-the-art.
関連論文
- DARP: 多視点ロボット知覚のための校正済み双腕RGB-D-IRデータセットデータセット
- uScenes: 水中ロボット知覚のためのマルチモーダルRGB・3Dソナー画像データセットデータセット
- PRISM:マルチモーダルセンシングを備えた精密で接触豊富な実世界産業スキルデータセットデータセット
- NARRATE: 自動運転における人間中心の説明のためのマルチモーダル実世界オーストラリア運転データセットデータセット
- 衛星画像の改ざんとディープフェイク位置特定のためのベンチマークデータセット構築に向けてデータセット
- InteracVid: ライブチャット動画から構築した実インタラクティブ音声視覚応答データセットデータセット