FUSEye: Training-Light Fisheye Detection with Overlapping Views and Zero-Initialized Adapters

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the severe performance degradation of COCO-pretrained detectors under fisheye distortion and the prohibitive cost of full fine-tuning by proposing an architecture-agnostic, lightweight adaptation framework. The method introduces a novel three-level causal linkage mechanism spanning input, feature, and decision stages. By integrating overlapping grid views, zero-initialized residual adapters, and consistent evidence fusion, it efficiently transforms a frozen-backbone YOLO into a robust fisheye detector. Evaluated on the WoodScape benchmark, the approach improves mAP50 from 0.148 to 0.266, retaining 84.3% of full fine-tuning accuracy with minimal additional parameters. Furthermore, it achieves 97.6% of its performance using only 25% of the training labels, significantly bridging the domain gap between standard and fisheye imagery.
📝 Abstract
Fisheye cameras give mobile robots a single-sensor, low-cost view of their surroundings, yet the COCO-pretrained detectors that practitioners routinely reuse fail on them: strong radial distortion warps local image structure, while boundary compression shrinks objects to near-invisible sizes. Full fine-tuning closes much of the gap but requires abundant fisheye labels and compute. We present FUSEye, a training-light framework that turns a frozen-backbone COCO-pretrained extra-large YOLO26 detector (YOLO26-x) into a fisheye detector. FUSEye adds roughly 227k new parameters while updating the inserted modules and the pretrained detection head. It addresses the transfer gap at three causally linked levels. At the input level, overlapping grid view generation and box remapping (GridViews) enlarge compressed boundary regions. At the feature level, zero-initialized residual adapters (Z-Adapters) correct distortion-induced feature misalignment. At the decision level, learned cross-projection agreement fusion (AgreeFusion) promotes low-confidence detections only when they are supported by consistent evidence across multiple views. On the WoodScape surround-view fisheye benchmark, FUSEye raises YOLO26-x from 0.148 to 0.266 mAP50 and retains 84.3% fully fine-tuned accuracy. Moreover, randomly using only 25% of the labeled training images, FUSEye achieves 0.2597 mAP50, retaining 97.6% of its full-label performance. FUSEye also consistently improves YOLOv8-11 detectors, showing that the recipe is architecture-agnostic. Source code will be available at https://github.com/Su-wenya/FUSEye.
Problem

Research questions and friction points this paper is trying to address.

fisheye detection
radial distortion
domain adaptation
training-light
object detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fisheye Detection
Training-Light Adaptation
Zero-Initialized Adapters
Overlapping Grid Views
Cross-Projection Fusion
🔎 Similar Papers
No similar papers found.