日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
魚眼検出/転移学習arXiv:2610.02799

FUSEye: 重複ビューとゼロ初期化アダプタによる学習軽量な魚眼検出

FUSEye: Training-Light Fisheye Detection with Overlapping Views and Zero-Initialized Adapters

シェア:XThreadsFacebookLINEはてブBluesky

凍結したCOCO事前学習済みYOLO検出器に約22.7万パラメータのアダプタを追加し、重複グリッドビューとゼロ初期化残差アダプタ、合意融合で魚眼画像の検出精度を大幅に向上させる学習軽量フレームワーク。

詳しい要約

1. どんなもの?

- モバイルロボット向けの魚眼カメラ検出器を、学習負荷を抑えて構築するフレームワーク FUSEye の提案。 - COCO 事前学習済みの extra-large YOLO26 (YOLO26-x) を backbone 凍結のまま魚眼検出器へ転換。 - 新規パラメータは約 227k のみで、挿入モジュールと検出 head を更新。 - 入力・特徴・決定の3段階で転移ギャップに対処する。

2. 先行研究と比べてどこがすごい?

- 実務で再利用される COCO 事前学習検出器は魚眼で失敗しやすい。 - 完全 fine-tuning は魚眼ラベルと計算資源を大量に要する。 - FUSEye は学習負荷を抑えつつ、WoodScape で YOLO26-x を 0.148 から 0.266 mAP50 へ改善。 - 完全 fine-tuning 精度の 84.3% を保持。 - ラベル付き学習画像をランダムに 25% のみ使用しても 0.2597 mAP50 を達成し、全ラベル性能の 97.6% を保持。 - YOLOv8-11 検出器も一貫して改善し、アーキテクチャ非依存なレシピであることを示す。

3. 技術・手法の肝は?

- 入力レベル: overlapping grid view generation と box remapping (GridViews) により、圧縮された境界領域を拡大。 - 特徴レベル: zero-initialized residual adapters (Z-Adapters) により、歪み由来の特徴ミスアライメントを補正。 - 決定レベル: learned cross-projection agreement fusion (AgreeFusion) により、複数ビューで一貫した証拠がある低信頼検出のみを促進。 - 3段階は因果的に関連付けられている。

4. どうやって有効だと検証した?

- WoodScape surround-view fisheye benchmark で評価。 - YOLO26-x を 0.148 から 0.266 mAP50 へ改善。 - 完全 fine-tuning 精度の 84.3% を保持。 - ラベル付き学習画像をランダムに 25% のみ使用した場合、0.2597 mAP50 を達成し、全ラベル性能の 97.6% を保持。 - YOLOv8-11 検出器でも一貫した改善を確認。

5. 議論はある?

- 要旨からは不明。 - 限界や失敗事例、計算コストの詳細な議論は要旨に記載されていない。

6. 次に読むべき論文は?

- YOLO26-x (YOLO26) および YOLOv8-11 検出器。 - COCO 事前学習検出器。 - WoodScape surround-view fisheye benchmark。 - 関連手法として GridViews、Z-Adapters、AgreeFusion。 - ソースコード: https://github.com/Su-wenya/FUSEye

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Wenya Su, Kai Luo, Di Wen, Ruiping Liu, Yufan Chen, Junwei Zheng, Kunyu Peng, Kailun Yang

分類: cs.CV, cs.RO, eess.IV

原文アブストラクト

Fisheye cameras give mobile robots a single-sensor, low-cost view of their surroundings, yet the COCO-pretrained detectors that practitioners routinely reuse fail on them: strong radial distortion warps local image structure, while boundary compression shrinks objects to near-invisible sizes. Full fine-tuning closes much of the gap but requires abundant fisheye labels and compute. We present FUSEye, a training-light framework that turns a frozen-backbone COCO-pretrained extra-large YOLO26 detector (YOLO26-x) into a fisheye detector. FUSEye adds roughly 227k new parameters while updating the inserted modules and the pretrained detection head. It addresses the transfer gap at three causally linked levels. At the input level, overlapping grid view generation and box remapping (GridViews) enlarge compressed boundary regions. At the feature level, zero-initialized residual adapters (Z-Adapters) correct distortion-induced feature misalignment. At the decision level, learned cross-projection agreement fusion (AgreeFusion) promotes low-confidence detections only when they are supported by consistent evidence across multiple views. On the WoodScape surround-view fisheye benchmark, FUSEye raises YOLO26-x from 0.148 to 0.266 mAP50 and retains 84.3% fully fine-tuned accuracy. Moreover, randomly using only 25% of the labeled training images, FUSEye achieves 0.2597 mAP50, retaining 97.6% of its full-label performance. FUSEye also consistently improves YOLOv8-11 detectors, showing that the recipe is architecture-agnostic. Source code will be available at https://github.com/Su-wenya/FUSEye.

PR本紙発行元 EmplifAI