Score
Designs and implements neural network modules (Feature Pyramid Networks, FPNs) that fuse, upsample, and combine multi-scale convolutional feature maps into a unified multi-level representation while preserving spatial detail across scales. These fused feature pyramids are analyzed and supplied to task-specific heads to improve localization and bounding-box prediction and to increase recall for small objects in image-based detection and segmentation.
To address the low feature representation ratio and weak spatial awareness of tiny objects in Feature Pyramid Networks (FPNs), this paper proposes a High-frequency and Spatial-aware FPN (HS-FPN). The method introduces two key components: (1) a High-frequency Perception (HFP) module that applies high-pass filtering to extract and enhance discriminative high-frequency features critical for tiny objects; and (2) a Spatial Dependency Perception (SDP) module that jointly models spatial-channel dependencies and cross-scale spatial relationships to strengthen both local and global spatial awareness. HS-FPN is an end-to-end trainable architecture requiring no post-processing. Evaluated on the AI-TOD benchmark, it achieves a significant mAP improvement over state-of-the-art methods, demonstrating the effectiveness of synergistically enhancing high-frequency content and modeling spatial structure for robust tiny-object representation learning.
To address computational redundancy and ambiguous boundary reconstruction in cross-layer feature pyramid networks (CFPNs) for salient object detection, this paper proposes the Context-aware Feature Mamba-based Dynamic fusion network (CFMD). Methodologically, CFMD introduces two key components: (1) a Context-aware Feature Long-range Memory Aggregation module (CFLMA) built upon the Mamba architecture, enabling efficient, long-range dependency modeling and dynamic weight allocation; and (2) an Adaptive Dynamic Upsampling unit (CFLMD) that combines bilinear initialization with a tunable receptive field mechanism to faithfully restore spatial details without degradation. Evaluated on three major benchmarks, CFMD achieves consistent improvements in both accuracy and efficiency: F-measure and E-measure increase significantly, boundary precision improves by 3.2% (Fβ) and 4.1% (Berkeley-DB), while inference speed rises by 18% (FPS), demonstrating superior trade-offs between real-time performance and pixel-level segmentation fidelity.
In multi-scale object detection, point-wise fusion in conventional feature pyramids causes misalignment between features across pyramid levels. Method: This paper proposes the Independent Hierarchical Pyramid (IHP) architecture, abandoning traditional top-down/bottom-up fusion paradigms. It introduces Soft Nearest-Neighbor Interpolation (SNI) and Extended Spatially Adaptive Downsampling (ESD) to establish an end-to-end secondary alignment (SA) mechanism—enabling precise, lightweight cross-scale feature matching without significant computational overhead. The method leverages GSConvE, a lightweight convolutional design, to enhance feature consistency without additional parameters. Contribution/Results: Evaluated on Pascal VOC and MS COCO, IHP achieves real-time state-of-the-art detection performance, with substantial AP gains for small objects. It jointly optimizes inference speed and accuracy, demonstrating superior efficiency–accuracy trade-offs.
Existing feature pyramid networks struggle to effectively model multi-scale discriminative features for dense visual prediction, particularly underperforming on small objects. This work proposes A3-FPN, a novel architecture that enables progressive decoupling for global feature interaction and incorporates a content-aware attention mechanism to enhance feature representation. During fusion and recombination stages, the method employs context-aware resampling and an information-driven redundancy optimization strategy, respectively, achieving efficient feature reassembly through positional offsets and content-adaptive weights. A3-FPN is compatible with both CNN and Transformer backbones and demonstrates significant performance gains across multiple benchmarks: it achieves 49.6 mask AP on MS COCO and 85.6 mIoU on Cityscapes when paired with OneFormer and Swin-L backbones, and also shows strong results on VisDrone2019-DET.
To address the semantic gap and information loss inherent in cross-layer interactions within Feature Pyramid Networks (FPNs) for small-object detection in aerial imagery, this paper proposes the Cross-layer Feature Pyramid Transformer (CFPT), an upsampling-free architecture. CFPT introduces two novel components: Cross-layer Channel/Space Attention (CCA/CSA) and Consistent Contextual Positional Encoding (CCPE), enabling linear-complexity, global-aware cross-layer feature fusion in a single step. Unlike conventional FPNs that rely on multi-stage upsampling and sequential top-down aggregation—exacerbating semantic mismatch—CFPT directly bridges hierarchical feature representations while preserving fine-grained spatial details critical for small objects. Evaluated on VisDrone2019-DET and TinyPerson benchmarks, CFPT achieves new state-of-the-art performance with lower computational overhead, improving mAP by 3.2% and 4.7%, respectively.
This work addresses the challenges of detecting rotated objects in high-resolution remote sensing imagery, where cluttered backgrounds, large scale variations, and complex orientations hinder performance. To tackle these issues, we propose a foreground-guided and angle-aware feature pyramid network that synergistically enhances object localization and orientation estimation. Specifically, foreground-guided feature modulation is introduced in low-level features to amplify responses in target regions, while an angle-aware multi-head attention mechanism is designed in high-level features to explicitly model directional geometric relationships. The model jointly optimizes foreground saliency and orientation priors under weak supervision. Our method achieves state-of-the-art results with mAP scores of 75.5% on DOTA v1.0 and 68.3% on DOTA v1.5, marking the first approach to successfully co-optimize foreground and angular information in a weakly supervised setting.
Existing object detectors often learn task-driven features that rely on shortcut correlations, failing to adequately capture the underlying annotation structure, which limits their generalization, interpretability, and robustness under task shifts or sparse supervision. To address this, this work proposes an annotation-guided feature enhancement framework that explicitly integrates geometric annotation priors into feature learning for the first time. By constructing a dense spatial feature grid and injecting it into the backbone network—where it fuses with the feature pyramid—the method steers region proposal and detection heads toward representations better aligned with annotation structure. Evaluated on wildlife and remote sensing datasets, the approach significantly improves object focus, reduces background sensitivity, and demonstrates superior generalization and data efficiency in weakly supervised and unseen-task settings.
This work addresses the challenges of salient object detection in optical remote sensing imagery, where large scale variations and complex backgrounds hinder existing methods from effectively modeling geometric structures and fine-grained details, often yielding incomplete detection results. To overcome these limitations, the authors propose G2HFNet, a novel architecture built upon the Swin Transformer backbone that introduces, for the first time, a geometry-granularity-aware mechanism. The model incorporates a hierarchical feature fusion framework comprising four key components: Multi-scale Detail Enhancement (MDE), Dual-branch Geometry-Granularity Complementary (DGC), Deep Semantic Perception (DSP), and Local-Global Guided Fusion (LGF). Extensive experiments demonstrate that G2HFNet significantly outperforms state-of-the-art methods across multiple remote sensing datasets, particularly excelling in complex scenes by producing more complete and accurate saliency maps.
This work addresses the challenges of object detection in drone imagery caused by complex background clutter and significant scale variation among targets. To this end, the authors propose SFFNet, a novel architecture that decouples multi-scale objects from the background through a frequency- and spatial-domain collaborative edge enhancement mechanism. The method innovatively integrates dual-domain dynamic edge extraction, linear deformable convolution, and a wide-range perception module, while introducing a collaborative feature pyramid network to strengthen both geometric and semantic representations. Additionally, a six-scale detection head is designed to enable precise localization. Evaluated on the VisDrone and UAVDT benchmarks, SFFNet achieves 36.8 AP and 20.6 AP, respectively, demonstrating substantial improvements in detection accuracy and parameter efficiency while maintaining a lightweight design.