Score
Designs and implements algorithms and modules that select and attend to spatial regions across multiple image scales, producing scale-aware spatial feature maps. Builds and analyzes mechanisms (e.g., visual adapters and dense multi-scale fusion) that choose scale-specific features and integrate them to improve localization and representation quality.
Existing responsive thematic mapping approaches rely heavily on manual intervention and employ disjointed visual encodings across devices, struggling to balance legibility with contextual consistency. This work proposes the first automated algorithmic framework that introduces a “layout-guided” structure to jointly model the visual requirements and relative spatial relationships of map elements. By integrating reference layouts, extremal horizontal and vertical order constraints, and rectangle- or Demers-based cartogram generation techniques, the map layout engine automatically produces stable and coherent responsive layouts tailored to container dimensions. The method enables smooth, deterministic, and visually consistent adaptation across arbitrary display sizes, significantly enhancing cross-device consistency and computational efficiency in thematic map generation.
This study addresses the challenge that existing automated map generalization methods struggle to jointly preserve spatial similarity and cartographic legibility across multiple scales, often treating similarity assessment, constraint modeling, and parameter optimization in isolation. To overcome this limitation, the authors propose a unified similarity-driven framework that formulates map generalization as a constrained multi-scale similarity optimization problem. For the first time, geometric, structural, and learned similarity measures are integrated into the objective function, while cartographic constraints—including legibility, smoothness, and geometric validity—are incorporated through line simplification algorithms. Experimental results demonstrate that the approach adaptively and consistently optimizes parameter configurations across diverse algorithms and scales, achieving high-quality map abstraction that maintains spatial similarity while significantly improving the interpretability and generalizability of parameter control.
To address the conflict between object-scale diversity in satellite imagery and constrained computational resources, this paper proposes a scale-adaptive recognition framework. Methodologically, it introduces large language models (LLMs) for the first time to perform semantic-driven scale concept reasoning; designs an active high-resolution (HR) image sampling strategy based on model disagreement; and establishes an HR–low-resolution (LR) cross-resolution knowledge distillation mechanism. The core innovation lies in jointly modeling LLM priors, model uncertainty, and multi-scale representations to enable budget-aware, on-demand HR invocation. Under stringent resource constraints, the framework achieves a 26.3% improvement in recognition accuracy over the full-HR baseline while reducing HR image usage by 76.3%, significantly enhancing deployment efficiency and cost-effectiveness.
To address global spatial layout inconsistency in high-resolution panoramic image generation, this paper proposes a multi-scale diffusion framework that jointly models geometry and semantics via cross-resolution structural prior transfer. The core contribution is a novel multi-scale gradient-guided mechanism, which explicitly injects low-resolution layout constraints into the high-resolution generation process, integrated with multi-scale feature distillation, cross-resolution gradient backpropagation, and structure-aware loss. The method preserves seamless structural coherence and natural transitions even at 16K resolution. Quantitative experiments demonstrate a 23.6% reduction in Fréchet Inception Distance (FID) and a 31.4% improvement in layout consistency metrics, significantly outperforming state-of-the-art approaches.
To address instance incompleteness and geometric shape mismatch in vectorized high-definition map construction from surround-view imagery—caused by frontal-view feature loss—this paper proposes a dual-path feature enhancement framework. First, we design a novel dual-enhancement module integrating explicit fusion and implicit modulation to enable efficient multi-view feature collaboration. Second, we introduce frontal-view keypoint supervision to guide BEV feature learning with geometric structural priors. Third, we develop an end-to-end BEV vectorization decoder. Evaluated on nuScenes and OpenLane benchmarks, our method achieves state-of-the-art performance, significantly improving instance completeness and geometric accuracy for lane markings, curbs, and other road elements—yielding +3.2% in F-score and +2.8% in AP₅₀.
This work addresses the limited generalization capability of object detection in aerial imagery caused by discrepancies in spatial resolution, scene composition, and semantic labeling across domains. To tackle this challenge, the authors propose a hierarchical modular routing framework that integrates geospatial-aware global routing with category-semantic-conditioned expert modules. By jointly leveraging global expert allocation and local scene decomposition, the method enables dual specialization—across datasets and within individual scenes—and supports zero-shot detection of novel categories without fine-tuning. Extensive experiments on four aerial image datasets demonstrate substantial improvements in multi-domain generalization, region-specific detection accuracy, and open-category recognition performance.
This work addresses the inefficiency of existing cross-view geolocalization methods in large-scale retrieval, their sensitivity to variations in satellite image resolution, and the error propagation inherent in hierarchical search strategies. To overcome these limitations, we propose GeoMoE, the first approach to introduce sparse Mixture-of-Experts (MoE) into this task. GeoMoE employs a dual-encoder architecture that decouples global multi-scale representation learning from local hierarchical search. By integrating content-adaptive routing and multi-scale supervision, it constructs a resolution-invariant embedding space, while probabilistic beam search substantially reduces computational cost and error accumulation. On the Just Zoom In benchmark, GeoMoE achieves 95.78% R@40m (+2.77%), and on the new VIGOR-M benchmark, it attains 62.39% R@1 with only 0.885 MMAC per query—just 5.27% of the cost of exhaustive L3 scanning—significantly outperforming the strongest baselines.
This work addresses the fundamental challenge in panoramic scene understanding: geometric distortions induced by spherical-to-planar projection violate the translation invariance assumption of conventional vision models, making it difficult to simultaneously achieve strict spherical equivariance and reuse pre-trained model weights. The study systematically establishes a two-dimensional taxonomy encompassing architectural design and training paradigms, revealing this core tension for the first time. It identifies five systemic deficiencies in current evaluation protocols and proposes a six-point research roadmap toward general-purpose panoramic intelligence. By integrating spherical geometric modeling, distortion-aware networks, native spherical operators, and geometry-aware tokenization, the authors construct a unified framework that pinpoints critical research gaps, advocates for standardized benchmarking, and releases an open-source survey repository to foster community progress.
Existing referring expression segmentation methods for remote sensing rely on full-parameter fine-tuning, which incurs high computational costs and risks degrading the generalization capability of pretrained models. Meanwhile, current parameter-efficient tuning approaches struggle to model cross-modal dependencies and bridge the domain gap between natural and aerial images. To address these limitations, this work proposes the S4ECA framework, featuring a dual-encoder adapter architecture: a text adapter generates high-level, learnable linguistic proxies, while a visual adapter integrates multi-scale dense features. A novel semantic-driven scale and spatial selection mechanism enables language-guided dynamic region focusing. By updating only 2.4% of the backbone parameters, S4ECA achieves state-of-the-art performance on both RRSIS-D and RefSegRS benchmarks, significantly enhancing segmentation accuracy and computational efficiency in complex aerial scenes.