Score
Designs and implements network components and modules that combine feature maps from multiple spatial resolutions, network depths, or sensor/channel streams into unified representations that retain fine-grained detail and contextual information. Builds and analyzes fusion blocks — e.g., pyramidal, cross-level, hierarchical or nested U‑Net style layers — that aggregate encoder/decoder maps, preserve boundaries and local texture, and compensate for missing or degraded inputs across scales.
Current visual systems exhibit insufficient robustness to noise, deformation, and out-of-distribution (OOD) data, and lack unsupervised, compositional, and interpretable neural representation mechanisms. To address these limitations, we propose the Collaborative Network Architecture (CNA), which introduces a novel “network fragment” dynamic composition mechanism. CNA unsupervisedly discovers local structural primitives via statistical learning and recursively assembles them through structured sparse connectivity, yielding global–local coupled representations that adaptively encode sensory patterns. Without any labeled data, CNA achieves robust recognition under noise and geometric deformation, zero-shot schema completion, and generalization to unseen patterns—significantly enhancing OOD generalization. Its modular, interpretable neural representations establish a new paradigm for invariant object recognition, bridging compositional structure with biological plausibility and computational efficiency.
This work addresses the limited generalization capability of object detection in aerial imagery caused by discrepancies in spatial resolution, scene composition, and semantic labeling across domains. To tackle this challenge, the authors propose a hierarchical modular routing framework that integrates geospatial-aware global routing with category-semantic-conditioned expert modules. By jointly leveraging global expert allocation and local scene decomposition, the method enables dual specialization—across datasets and within individual scenes—and supports zero-shot detection of novel categories without fine-tuning. Extensive experiments on four aerial image datasets demonstrate substantial improvements in multi-domain generalization, region-specific detection accuracy, and open-category recognition performance.
Existing node embedding methods suffer from two key limitations: vector addition lacks network semantic interpretability, and relationships among multi-scale (coarse-grained) embeddings remain ill-defined. This paper proposes a multi-scale node embedding framework that unifies the resolution of both issues for the first time. Leveraging a hierarchical coarse-graining mechanism grounded in renormalization theory, and imposing vector-sum constraints alongside low-dimensional reconstruction optimization in the embedding space, our method ensures that the embedding of any coarse-grained block node is strictly equal to the statistical mean of its constituent node embeddings. This guarantees statistical consistency across resolutions. Evaluated on international trade and input-output networks, the framework achieves high-fidelity structural reconstruction—e.g., accurate triangle counting—and supports arbitrary-scale graph generation. It significantly enhances interpretability and practicality in multi-scale graph modeling and synthesis.
Existing neural networks are constrained by hierarchical tree-like architectures, which preclude direct communication among sibling nodes and prohibit backward signal propagation to higher-level modules—resulting in weak inter-module collaboration and inefficient representation learning. To address these limitations, we propose the Synchronous Graph Neural Architecture (SGNA), organizing neural units into a modular, dynamically collaborative graph structure that enables arbitrary node-to-node communication and cross-layer signal transmission. Our key contributions are threefold: (1) introducing the first modular graph-structured paradigm for neural architecture design; (2) developing a systematic regularization framework to enforce module independence and load balancing; and (3) generalizing neural architecture search (NAS) to the space of directed acyclic graphs (DAGs). Extensive multi-task experiments demonstrate that SGNA significantly outperforms deep stacked baselines under parameter constraints, achieving superior collaborative representation capability and more comprehensive coverage of the search space.
To address instance incompleteness and geometric shape mismatch in vectorized high-definition map construction from surround-view imagery—caused by frontal-view feature loss—this paper proposes a dual-path feature enhancement framework. First, we design a novel dual-enhancement module integrating explicit fusion and implicit modulation to enable efficient multi-view feature collaboration. Second, we introduce frontal-view keypoint supervision to guide BEV feature learning with geometric structural priors. Third, we develop an end-to-end BEV vectorization decoder. Evaluated on nuScenes and OpenLane benchmarks, our method achieves state-of-the-art performance, significantly improving instance completeness and geometric accuracy for lane markings, curbs, and other road elements—yielding +3.2% in F-score and +2.8% in AP₅₀.
To address two key bottlenecks in camouflaged object detection (COD) and salient object detection (SOD)—insufficient intra-layer channel-wise interaction and difficulty in jointly modeling boundary and region information—this paper proposes a Channel Information Interaction Module (CIIM) and a prior-guided collaborative decoding architecture. CIIM enables cross-channel feature reorganization via horizontal–vertical channel integration, while the decoder employs dual-path prior generation (boundary/region) coupled with hybrid attention-based calibration to jointly optimize structural and semantic cues. Notably, this is the first work to unify COD and SOD under a single framework, demonstrating strong cross-task generalization. Our method achieves state-of-the-art performance on four COD benchmarks. Moreover, it successfully transfers to diverse downstream tasks—including SOD, polyp segmentation, transparent object detection, and industrial defect detection—validating its robustness and versatility. Code and comprehensive experimental results are publicly available.
This study addresses the design of efficient, general-purpose, and scalable embedding representations for Earth observation (EO) tasks by systematically evaluating key architectural and training choices when using GeoFM as a feature extractor. These include backbone architecture, pretraining strategy, representation depth, spatial aggregation, and multi-objective fusion. Experiments on the NeuCo-Bench benchmark demonstrate that a Transformer backbone combined with mean pooling establishes a strong baseline; intermediate ResNet layers outperform final-layer features; self-supervised objectives offer task-specific advantages; and multi-objective fusion substantially enhances robustness. The resulting embeddings compress the original input by over 500× while maintaining high performance across diverse downstream EO tasks, underscoring the critical role of thoughtful embedding design in enabling scalable EO workflows.
To address insufficient coordination between grouping and feature extraction layers in point cloud networks—which hinders full exploitation of raw point potentials—this paper proposes a module-level optimization approach. We design a lightweight Grouping-Feature Coordination (GF-Core) module, the first to jointly and dynamically regulate both grouping and feature extraction layers. To enhance geometric fidelity and discriminability, we introduce an attention-guided separable mechanism and a coordinate-feature joint similarity-based grouping strategy. Furthermore, we develop a self-supervised contrastive pretraining framework tailored for point clouds to improve robustness. Our method is architecture-agnostic and achieves 94.0% accuracy on ModelNet40—comparable to state-of-the-art methods—while outperforming baselines by 2.96%, 6.34%, and 6.32% on the three ScanObjectNN variants, respectively. These results demonstrate significant gains in generalization and robustness under real-world scanning conditions.
This study addresses the challenge of fusing multi-source nighttime light data (DMSP-OLS and VIIRS) in satellite remote sensing, where conventional methods suffer from limited spatial and temporal resolution. We systematically explore the design space of diffusion models and normalizing flow models for image-level data fusion. We propose a UNet-based diffusion model featuring a novel staged noise scheduling strategy and mixed-precision quantization, jointly optimizing generation fidelity, fine-detail preservation, inference efficiency, and computational deployability. Experiments demonstrate superior performance over traditional and state-of-the-art generative models on cross-sensor and cross-spatiotemporal-resolution fusion tasks, achieving a PSNR improvement of over 2.1 dB. The work establishes a reproducible generative modeling paradigm for remote sensing data fusion, provides principled architectural design guidelines, and delivers a lightweight optimization pathway for practical deployment.
Existing camouflaged object detection methods suffer from computational redundancy due to stacked boundary modules and attention mechanisms, and often operate at low resolutions, leading to loss of critical edge details, edge discontinuities, and fragmentation. To address these issues, we propose a Collaborative Perception-Guided Unified Network (CPG-UNet). It jointly models contextual information via channel calibration and spatial enhancement; introduces multi-scale feature fusion and progressive boundary optimization to achieve scale-adaptive edge modulation at intermediate resolution—balancing semantic consistency and fine-grained detail preservation. Extensive experiments demonstrate state-of-the-art performance: Sα scores of 0.887, 0.890, and 0.895 on CAMO, COD10K, and NC4K benchmarks, respectively. Our method significantly improves detection accuracy for small objects, large background-similar objects, and heavily occluded scenarios, while maintaining real-time inference capability (≈32 FPS on a single RTX 3090).