spatial prior integration

Designs and implements mechanisms to condition, fuse, or inject external spatial priors into machine-learning models, creating modules or pipelines that accept anatomical, structural, motion, land-cover, or other spatial maps and use them to influence generation, denoising, or feature representations. These implementations represent priors as activation maps or attention masks, aggregate or fuse them with model features (via conditioning, auxiliary networks, or feature-level fusion), and are evaluated by their ability to anchor outputs to spatial cues, stabilize boundaries, disambiguate ambiguous structures, and improve downstream metrics such as segmentation, reconstruction, or motion estimates.

spatialpriorintegration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.29
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

In complex urban driving scenarios lacking high-definition maps, existing methods suffer from insufficient utilization of structural road priors, resulting in irregular and brittle predictions. To address this, we propose a unified road perception framework that jointly integrates semantic, geometric, and generative priors. Our approach introduces three key innovations: (1) an instance-aware, shape-guided attention mechanism that explicitly models geometric structure during feature aggregation; (2) a data-driven shape template space constructed via clustering, enabling compact, low-dimensional anchor priors for robust road element representation; and (3) a diffusion-based structured generation paradigm that enforces geometric consistency and regularity in output predictions. Evaluated on large-scale autonomous driving benchmarks, our method achieves significant improvements in road element detection accuracy. Notably, it demonstrates superior robustness and prediction consistency under challenging conditions—including severe occlusion and rapid curvature variation—while maintaining strong generalization across diverse urban layouts.

Enhancing road perception in autonomous driving without HD mapsIntegrating semantic, geometric, and generative priors for accuracyOvercoming occlusion and complexity challenges in road elements

FeatInv: Spatially resolved mapping from feature space to input space using conditional diffusion models

May 27, 2025
NN
Nils Neukirch
🏛️ Carl von Ossietzky Universität Oldenburg | Fraunhofer Heinrich-Hertz-Institute

Deep neural network feature spaces lack interpretability, particularly due to the absence of an exact, pixel-level inverse mapping from spatial feature maps back to input images. To address this, we propose Spatially Aligned Conditional Diffusion (SACD), the first method to condition high-fidelity pre-trained diffusion models—pixel-wise—on spatially resolved feature maps extracted from CNNs or Vision Transformers. SACD enables a probabilistic, high-fidelity, and sampleable inverse mapping from features to inputs. Crucially, it supports concept-guided visualization (“concept steering”) and feature disentanglement analysis without fine-tuning or restrictive approximations. We evaluate SACD on multiple ImageNet-pretrained classifiers, demonstrating superior reconstruction quality and robustness over existing feature inversion approaches. Our method establishes a new paradigm for interpretable deep representation learning by bridging generative modeling and feature-space analysis in a principled, scalable manner.

Improving understanding of feature space in vision modelsMapping feature space to input space for interpretabilityUsing conditional diffusion models for high-fidelity reconstruction

Existing multimodal large language models rely on a single visual prior, limiting their ability to effectively support diverse spatial understanding tasks. To address this, this work proposes the ViPS framework, which systematically reveals—for the first time—the complementary nature of multiple visual priors. ViPS introduces lightweight prior proxies and a dynamic fusion mechanism that enables context-aware collaborative injection of heterogeneous priors. By moving beyond the constraints of a single expert representation, the method achieves significant performance gains over current state-of-the-art approaches across multiple challenging benchmarks involving complex spatial reasoning and 3D understanding, establishing new best results.

Foundation ModelsMultimodal Large Language ModelsPrior Integration

This study investigates the true role and underlying mechanism of positional encoding in Vision Transformers (ViTs) with respect to spatial geometric reasoning. Addressing the lack of deep understanding regarding the geometric significance of positional encoding in existing literature, we propose a token-level multi-view geometric consistency diagnostic framework, which for the first time demonstrates that positional encoding acts as a causal factor in shaping the spatial structure of ViT representations. Through comprehensive ablation and probing experiments across 14 foundational ViT models, we validate that positional encoding simultaneously guides both local structural coherence and global layout organization, thereby establishing its essential role as a critical geometric prior in ViT architectures.

geometric priorsmulti-view geometrypositional embeddings

Spatial Reasoning with Denoising Models

Feb 28, 2025
CW
Christopher Wewer
🏛️ Max Planck Institute for Informatics

Existing spatial generative models—such as diffusion models and flow matching—suffer from hallucination when modeling complex distributions, leading to distorted inference and mode collapse. To address this, we propose the Spatial Reasoning Model (SRM) framework, the first to adapt denoising-based generation for causal or constraint-driven reasoning over continuous variable sets, enabling high-fidelity continuous inference from observed to unobserved variables. Our key contributions are: (1) a learnable, adaptive generation-order prediction mechanism; (2) a continuous reasoning architecture explicitly designed for spatial constraints; and (3) the first benchmark suite quantifying generative hallucination in spatial reasoning. Experiments demonstrate that SRM boosts accuracy on targeted spatial reasoning tasks from <1% to >50%, substantially mitigates hallucination, and empirically validates both the learnability of generation order and its decisive impact on reasoning performance.

Addresses hallucination in generative models for complex distributions.Demonstrates improved reasoning accuracy via denoising network predictions.Introduces Spatial Reasoning Models for continuous variable reasoning.

Latest Papers

What's happening recently
View more

High-definition maps are costly to produce, and existing 3D detection methods struggle to effectively leverage environmental structural priors to address challenges such as sparse and noisy sensor data or adverse weather conditions. This work proposes the MPA3D framework, which automatically constructs dense, annotation-free scene reconstructions from multi-frame LiDAR and image data to serve as mapping priors. These priors are systematically integrated into an end-to-end 3D object detection pipeline, enabling deep fusion of multimodal features and structural context. Evaluated on the Waymo Open Dataset, the approach achieves state-of-the-art performance, demonstrating that scalable reconstruction priors provide significant benefits for 3D detection accuracy and robustness.

3D object detectionautonomous drivingHD maps

Existing spherical-based fMRI decoders struggle to balance performance and geometric fidelity due to inefficient spherical tokenization and the neglect of individual anatomical structure. This work addresses these limitations by explicitly modeling cortical anatomical features as inductive priors, introducing a Selective Region-of-Interest Spherical Tokenizer (SRST) and a Structure-Guided Mixture-of-Experts module (SG-MoE) to enable efficient geometric encoding and seamless integration of anatomical information. Evaluated on the Natural Scenes Dataset, the proposed method achieves a new state-of-the-art in surface-based decoding, matching the performance of leading 1D baselines while accelerating training convergence by 30× and enabling rapid adaptation to new subjects with only 20% of the data.

anatomical priorscortical anatomyfMRI decoding

This work addresses the limitation of existing end-to-end autonomous driving approaches, which rely solely on instantaneous sensor inputs and lack the human driver’s experience-based anticipatory capability. To overcome this, the authors propose a memory-augmented architecture that integrates geospatial visual priors through dual memory modules and an adaptive memory gating mechanism. Guided by planned trajectories, the system enables spatial anticipation without dependence on real-time perception while retaining a safety fallback capability. Evaluated on the NAVSIM-v2 benchmark, the method significantly enhances the performance of multiple end-to-end baselines and demonstrates robust, reliable driving behavior even under sensor failure conditions, thereby improving overall system robustness.

anticipatory behaviorautonomous drivinggeospatial priors

The geometric structure of intermediate feature spaces in deep neural networks remains poorly understood. This work systematically investigates the mapping between original and transformed features by applying geometric and photometric transformations, local masking, and generative semantic edits to input images. The experiments reveal that, to a first-order approximation, the feature space exhibits a linear structure: a shared linear model can efficiently and faithfully implement diverse complex semantic operations, matching or significantly outperforming more sophisticated nonlinear and Transformer-based models. These findings challenge the prevailing reliance on nonlinear modeling in representation learning and provide crucial empirical evidence toward understanding the intrinsic geometric nature of deep network feature spaces.

deep neural networksfeature representationsfeature space

This work addresses the limitations of existing object placement methods, which rely on manual annotations or artifact-prone inpainting pipelines, leading to poor scalability and susceptibility to shortcut learning. The authors propose a fully automatic and scalable evaluation framework that, for the first time, efficiently distills a generalizable class-conditional spatial prior from text-to-image diffusion models. They construct HiddenObjects, a large-scale dataset comprising 27 million annotations, and introduce a diffusion-based image inpainting strategy coupled with a dense object insertion evaluation protocol to compress this prior into a lightweight inference model. On downstream image editing tasks, the method substantially outperforms sparse human annotations (3.90 vs. 2.68 VLM-Judge) and achieves a 230,000× speedup at inference time, significantly surpassing current baselines and zero-shot vision-language models.

diffusion modelsnatural scenesobject placement

Hot Scholars

UB

Ulas Bagci

Northwestern University
artificial intelligencedeep learningbiomedical image analysismedical image computing
MP

Marc Pollefeys

Professor of Computer Science, ETH Zurich, and Director Spatial AI Lab, Microsoft
Computer VisionComputer GraphicsRoboticsMachine Learning
GD

Gorkem Durak

Northwestern University, Department of Radiology
radiologyartificial intelligence
EK

Elif Keles

Northwestern University
pediatricsneuroscienceneonatologyartificial intelligence
YZ

Yefeng Zheng

Professor, Westlake University, Hangzhou, China, IEEE Fellow, AIMBE Fellow
AI in HealthMedical ImagingComputer VisionNatural Language Processing