angular-modulated spatial fusion

Designs and implements neural modules and transformer architectures that fuse spatial image features with angular (viewpoint) information, by modulating spatial feature fusion with angular priors and refining inter-view interactions. These components analyze and enforce geometry-aware coherence across views or rays to reduce view-dependent artifacts and preserve 4D ray consistency during feature aggregation and reconstruction.

angular-modulatedspatialfusion

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.4
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses a key limitation in existing feed-forward neural view synthesis (NVS) Transformers, where semantic and spatial information are entangled within a shared feature space, causing spatial bias to interfere with appearance representation and degrade rendering fidelity. To resolve this, the authors propose a semantics-spatial disentangled architecture that explicitly separates feature representations into independent branches while enabling efficient cross-branch interaction through shared attention routing. Additionally, they introduce optional classification supervision and a bidirectional modulation mechanism to enhance representational capacity with negligible impact on inference latency. The proposed approach consistently improves performance across both decoder-only and encoder-decoder variants of feed-forward NVS models, yielding significantly higher rendering quality.

novel view synthesisrendering fidelityrepresentation ambiguity

AngularFuse: A Closer Look at Angle-based Perception for Spatial-Sensitive Multi-Modality Image Fusion

Oct 14, 2025
XL
Xiaopeng Liu
🏛️ Guangdong University of Technology | Sun Yat-sen University | TikTok | ByteDance Inc | Anhui University

Existing unsupervised visible-infrared image fusion methods suffer from handcrafted loss functions: reference images often lack fine details and exhibit uneven brightness, while gradient losses model only magnitude—ignoring directional information—leading to spatial structural distortion. To address these issues, we propose AngularFuse, the first angle-aware fusion framework that jointly constrains both gradient magnitude and direction in the gradient domain. We design a cross-modal complementary masking module and integrate Laplacian edge enhancement with adaptive histogram equalization to collaboratively generate high-quality reference images. Extensive experiments on MSRS, RoadScene, and M3FD datasets demonstrate that AngularFuse significantly outperforms state-of-the-art methods, especially under low-light conditions and complex backgrounds. It achieves superior detail fidelity and edge accuracy, thereby enhancing robustness and applicability for downstream vision tasks.

Addresses limitations in visible-infrared image fusion methodsEnhances texture intensity and edge orientation in fused imagesProposes angle-based perception for spatial-sensitive multimodal fusion

This study investigates the true role and underlying mechanism of positional encoding in Vision Transformers (ViTs) with respect to spatial geometric reasoning. Addressing the lack of deep understanding regarding the geometric significance of positional encoding in existing literature, we propose a token-level multi-view geometric consistency diagnostic framework, which for the first time demonstrates that positional encoding acts as a causal factor in shaping the spatial structure of ViT representations. Through comprehensive ablation and probing experiments across 14 foundational ViT models, we validate that positional encoding simultaneously guides both local structural coherence and global layout organization, thereby establishing its essential role as a critical geometric prior in ViT architectures.

geometric priorsmulti-view geometrypositional embeddings

This study addresses the challenge that the effects of module-level features and their interactions on generalization in Vision Transformer (ViT) architectures remain insufficiently isolated and quantified. By analyzing ViT representational structures, we identify feature collapse during initialization and propose quantitative metrics based on feature entropy and minimum eigenvalues. Our analysis reveals the critical roles of token spaces and linear submodules in generalization, systematically validated through multi-scale architectural comparison experiments. Results demonstrate that the proposed surrogate metrics improve correlation rankings with generalization performance by 18%–48%, enabling precise identification of low-compute, high-accuracy ViT architectures. This work provides both theoretical foundations and practical guidance for efficient visual model design.

Architecture DesignFeature InformationGeneralization Behavior

Understanding Multi-View Transformers

Oct 28, 2025
MS
Michal Stary
🏛️ TUM | Claude Bernard University Lyon 1 | University of Cambridge | MIT

Multi-view Transformers (e.g., DUSt3R) achieve strong performance in 3D vision but remain poorly understood, hindering interpretability and safe deployment. To address this, we propose the first layer-wise interpretability analysis framework for multi-view Transformers, leveraging residual connection feature probing and geometry-aware visualization to systematically uncover the evolution of internal 3D representations. Our analysis reveals that hidden states progressively encode 3D structure: early layers focus on local correspondences, while deeper layers refine these via geometric reconstruction—enabling end-to-end pose estimation without explicit global pose modeling. This work is the first to elucidate *how* and *why* such models succeed, breaking the “black-box” barrier and providing theoretical foundations for architecture design and reliability validation. Code is publicly available.

Analyzing inner mechanisms of multi-view transformers like DUSt3RInvestigating latent state development and layer roles in transformersVisualizing 3D representations from transformer residual connections

Latest Papers

What's happening recently
View more

Existing light field super-resolution methods struggle to effectively model spatial-angular correlations while preserving 4D ray consistency, primarily due to dimension-decoupled strategies and scanning mechanisms that violate epipolar geometry. To address these limitations, this work proposes SMART, a hybrid network that uniquely integrates angular priors into spatial modeling. SMART synergistically combines a slope-guided Mamba module with an angle-optimized Transformer, complemented by an angle-modulated spatial module and a manifold alignment trajectory module that adheres to epipolar geometry. This architecture enables geometrically consistent cross-dimensional feature learning. Evaluated on five benchmarks, the proposed method achieves state-of-the-art performance, yielding an average PSNR gain of 0.42 dB and significantly suppressing visual artifacts.

4D Ray CoherenceEpipolar GeometryLight Field Super-Resolution

This work addresses the challenges of view inconsistency and geometric distortion in single-image 3D reconstruction, which often arise from projection ambiguities during multi-view synthesis. To mitigate these issues, the authors propose a view-adaptive neural rendering framework that employs a shared feature backbone to capture global structure while enabling per-view independent correction of rendering errors. A lightweight self-attention fusion module is introduced to integrate multi-view information and enhance geometric consistency without relying on supervision from diffusion models such as SDS. The method optimizes solely with photometric loss, achieving near state-of-the-art reconstruction fidelity while maintaining computational efficiency and significantly improving view consistency and practical performance.

3D reconstructionmonocular 3D generationNeural Radiance Field

This study addresses the lack of systematic analysis regarding the geometric evolution of internal representations during Vision Transformer (ViT) training. The authors propose the TGO-II framework, which integrates Centered Kernel Alignment (CKA), Singular Vector Canonical Correlation Analysis (SVCCA), TwoNN intrinsic dimension estimation, and token covariance analysis. Applying this framework to ViT-Small/16 under supervised training, they uncover a tripartite geometric evolution pattern: progressive layer-wise specialization, an initial rise followed by stabilization of intrinsic dimensionality, and the persistent presence of strong token interaction structures. These findings demonstrate that increased representational complexity co-occurs with layer specialization without requiring token decorrelation, thereby challenging the conventional assumption that complexity arises from token independence and highlighting ViT’s capacity to achieve rich representational transformations through sustained strong token interactions.

intrinsic dimensionalityrepresentation geometryrepresentational similarity

This work addresses the scalability bottleneck in feed-forward novel view synthesis, where Transformer-based architectures suffer from prohibitive computational and memory overheads as the number of input views increases. To this end, we propose the AIMS framework, whose core innovation lies in decoupling the number of observations from the number of processed views via an anchor integration mechanism. Specifically, AIMS employs farthest point sampling to select a fixed set of anchors, combined with spatial clustering and a lightweight learnable integrator to aggregate information from neighboring observations. This design enables the efficient utilization of large-scale multi-view data under a constant global view budget. Experimental results demonstrate that AIMS achieves PSNR values of 29.41 dB on RealEstate10K and 17.73 dB on ScanNet, with an average rendering time of only 7.24 milliseconds per view, yielding a superior quality-efficiency trade-off.

Computational EfficiencyFeed-forwardMulti-View

Hot Scholars

NA

Nikos A. Mitsiou

PhD student, Aristotle University of Thessaloniki
Wireless CommunicationsOptimizationMachine Learning
HL

Hyeongtaek Lee

Assistant Professor of Ewha Womans University, Dept. of Electronic and Electrical Engineering
Wireless CommunicationsMassive MIMOReconfigurable Intelligent Surface
IK

Ioannis Krikidis

Professor, Electrical and Computer Engineering, IRIDA, University of Cyprus
Wireless CommunicationsMIMOWireless Power TransferNetworks
JC

Junil Choi

KAIST Endowed Chair Associate Professor, School of Electrical Engineering
Wireless CommunicationsSignal ProcessingMIMO
GL

Gyoseung Lee

School of Electrical Engineering, KAIST
Wireless CommunicationsSignal ProcessingReconfigurable Intelligent Surface