Score
Design and implement algorithms, modules, and pipelines that aggregate local measurements across spatial locations and scales into hierarchical, multi-scale descriptors and region-level summaries (e.g., spatial-pyramid pooling, multi-scale pooling, and other spatial aggregation functions). Analyze and evaluate these aggregation schemes to produce compact, semantically coherent hierarchical representations and to improve region-level correspondence or matching across scales.
Point cloud neighborhood aggregation suffers from interference by irrelevant points and feature-level disconnection; existing geometry-encoding methods incur high computational overhead and exhibit poor noise robustness. To address this, we propose a cross-stage structural correlation-driven neighbor aggregation correction framework. Specifically, we design a Point Distribution Set Abstraction (PDSA) module to model distributional correlations in high-dimensional space for aggregation refinement. We further introduce a lightweight cross-stage structural descriptor coupled with a keypoint mechanism to enhance structural homogeneity while reducing computational cost. Our method integrates neighborhood variance suppression, inter-class separability enhancement, and optimized keypoint sampling. Extensive experiments demonstrate that it significantly outperforms state-of-the-art baselines on semantic segmentation and classification tasks, achieving superior generalization with fewer parameters. Ablation studies and visualizations validate both its effectiveness and mechanistic soundness.
This work addresses subdomain aggregation over time-series and image data, proposing the first unified modeling framework grounded in category theory. Methodologically, it formalizes classical aggregation operations—including summation, extremum computation, and sliding-window statistics—as bifunctors on double categories, thereby achieving functional abstraction of aggregation semantics across diverse data structures. Integrating functorial semantic modeling with Blelloch’s parallel scan algorithm, the framework derives novel aggregation operators with provably parallel implementations, substantially extending the applicability of the scan paradigm. Key contributions are: (1) the first functional aggregation framework supporting formal verification of parallelizability; (2) systematic definition and generation of previously unformalized subdomain aggregation patterns; and (3) a composable, extensible mathematical foundation for cross-modal data aggregation. The approach bridges abstract categorical semantics with practical parallel computation, enabling rigorous, scalable, and interoperable aggregation across heterogeneous data modalities.
To address the high computational cost and model complexity in point cloud 3D detection caused by frequent neighborhood searches in multi-scale feature learning, this paper proposes a single-neighborhood multi-scale feature approximation framework. Our method avoids redundant neighborhood construction via knowledge distillation to compress and approximate multi-scale features; introduces a transferable, class-aware statistical embedding mechanism that leverages lightweight class-level statistics to compensate for diversity loss under single-neighborhood sampling; and adopts a center-weighted IoU loss to mitigate localization misalignment induced by center offset. Evaluated on Waymo and KITTI benchmarks, our approach achieves state-of-the-art detection accuracy while reducing inference latency by 2.1× and model parameters by 37%, demonstrating significant efficiency gains without compromising detection performance.
This paper addresses the challenge of automating choreme abstraction for categorical regional maps (e.g., isarithmic or land-use maps), where existing methods lack scalability and formal optimization. We propose a disk-based algorithmic summarization framework that models visual abstraction as a categorical point-set covering problem. By combining spatial sampling with combinatorial optimization, we efficiently compute a minimal set of representative disks capturing dominant spatial patterns. Our approach introduces the first extensible approximate algorithm framework for choreme generation, balancing visual fidelity and computational efficiency. Experiments demonstrate high-quality choreme summaries generated in seconds under moderate sampling densities; comparative evaluation across multiple sampling strategies confirms the method’s effectiveness, robustness, and practical applicability. This work bridges a critical gap in automated choreme synthesis for categorical maps and establishes a novel paradigm for geographic visualization summarization.
This work addresses the inefficiency of existing cross-view geolocalization methods in large-scale retrieval, their sensitivity to variations in satellite image resolution, and the error propagation inherent in hierarchical search strategies. To overcome these limitations, we propose GeoMoE, the first approach to introduce sparse Mixture-of-Experts (MoE) into this task. GeoMoE employs a dual-encoder architecture that decouples global multi-scale representation learning from local hierarchical search. By integrating content-adaptive routing and multi-scale supervision, it constructs a resolution-invariant embedding space, while probabilistic beam search substantially reduces computational cost and error accumulation. On the Just Zoom In benchmark, GeoMoE achieves 95.78% R@40m (+2.77%), and on the new VIGOR-M benchmark, it attains 62.39% R@1 with only 0.885 MMAC per query—just 5.27% of the cost of exhaustive L3 scanning—significantly outperforming the strongest baselines.
This work addresses the limitations of existing Earth observation foundation models, which predominantly rely on raster data and overlook the structured geographic semantics embedded in open vector datasets such as OpenStreetMap, thereby hindering comprehensive understanding of human–environment systems. To overcome this, we propose the first unified spatial representation learning framework that deeply integrates remote sensing imagery and vector data within a shared embedding space, breaking away from conventional modality-isolated paradigms. By leveraging self-supervised learning and multimodal alignment—while explicitly modeling geometric, topological, and semantic relationships—our approach enables synergistic raster perception and vector-based reasoning. The method substantially enhances accuracy, semantic interpretability, and explainability on downstream tasks, laying a theoretical and methodological foundation for developing human-centered, semantically rich geospatial foundation models.
This work addresses the challenge of extracting multi-granular, hierarchical 3D scene groupings from 2D foundation models without relying on semantic labels or a fixed vocabulary. To this end, it proposes transforming 2D affinity cues into hierarchical supervision signals and embedding them into a unified feature field within Lorentzian hyperbolic space. The method uniquely integrates the Dasgupta hierarchical clustering objective with hyperbolic geometry to construct a hyperspherical affinity field capable of representing groupings at multiple granularities. A lowest common ancestor ordering constraint is further introduced to enhance hierarchical consistency. This framework enables end-to-end, unsupervised 3D hierarchical grouping, accurately recovering object- and part-level structures across multiple scales while effectively fusing 2D priors with 3D geometric representations.
Traditional marked point process analyses rely on global statistics and struggle to capture local spatial heterogeneity. This study proposes the Local Indicator of Mark Association (LIMA) framework, which for the first time incorporates compositional marks into local spatial analysis. By leveraging the centered log-ratio (clr) transformation and Aitchison geometry, LIMA maps compositional data into Euclidean space, enabling a pointwise decomposition of mark structure. The method effectively uncovers local clustering and “siphoning” effects that are obscured by global approaches. Simulation experiments demonstrate that LIMA substantially outperforms existing global methods in detecting local clusters. When applied to economic data from Castilla–La Mancha, Spain, LIMA successfully reveals latent regional economic agglomeration patterns.