Score
Designs and implements neural modules and transformer architectures that fuse spatial image features with angular (viewpoint) information, by modulating spatial feature fusion with angular priors and refining inter-view interactions. These components analyze and enforce geometry-aware coherence across views or rays to reduce view-dependent artifacts and preserve 4D ray consistency during feature aggregation and reconstruction.
This paper addresses the fragmented and unsystematic modeling of geometric constraints in deep learning by proposing the first unified taxonomy of geometric constraints tailored for modern deep learning frameworks. Methodologically, it systematically integrates multi-view geometry, epipolar constraints, camera calibration models, self-supervised geometric consistency losses, and differentiable rendering to establish a three-dimensional classification framework spanning modeling principles, integration strategies, and optimization objectives. The contributions are threefold: (1) clarifying the applicability boundaries and failure mechanisms of over one hundred geometric constraints across vision tasks such as depth estimation; (2) uncovering key design paradigms for synergistic co-design of geometric priors and neural architectures; and (3) identifying principled pathways to overcome three core challenges—dynamic scenes, textureless regions, and cross-domain generalization.
This work addresses a key limitation in existing feed-forward neural view synthesis (NVS) Transformers, where semantic and spatial information are entangled within a shared feature space, causing spatial bias to interfere with appearance representation and degrade rendering fidelity. To resolve this, the authors propose a semantics-spatial disentangled architecture that explicitly separates feature representations into independent branches while enabling efficient cross-branch interaction through shared attention routing. Additionally, they introduce optional classification supervision and a bidirectional modulation mechanism to enhance representational capacity with negligible impact on inference latency. The proposed approach consistently improves performance across both decoder-only and encoder-decoder variants of feed-forward NVS models, yielding significantly higher rendering quality.
Existing unsupervised visible-infrared image fusion methods suffer from handcrafted loss functions: reference images often lack fine details and exhibit uneven brightness, while gradient losses model only magnitude—ignoring directional information—leading to spatial structural distortion. To address these issues, we propose AngularFuse, the first angle-aware fusion framework that jointly constrains both gradient magnitude and direction in the gradient domain. We design a cross-modal complementary masking module and integrate Laplacian edge enhancement with adaptive histogram equalization to collaboratively generate high-quality reference images. Extensive experiments on MSRS, RoadScene, and M3FD datasets demonstrate that AngularFuse significantly outperforms state-of-the-art methods, especially under low-light conditions and complex backgrounds. It achieves superior detail fidelity and edge accuracy, thereby enhancing robustness and applicability for downstream vision tasks.
This study investigates the true role and underlying mechanism of positional encoding in Vision Transformers (ViTs) with respect to spatial geometric reasoning. Addressing the lack of deep understanding regarding the geometric significance of positional encoding in existing literature, we propose a token-level multi-view geometric consistency diagnostic framework, which for the first time demonstrates that positional encoding acts as a causal factor in shaping the spatial structure of ViT representations. Through comprehensive ablation and probing experiments across 14 foundational ViT models, we validate that positional encoding simultaneously guides both local structural coherence and global layout organization, thereby establishing its essential role as a critical geometric prior in ViT architectures.
This study addresses the challenge that the effects of module-level features and their interactions on generalization in Vision Transformer (ViT) architectures remain insufficiently isolated and quantified. By analyzing ViT representational structures, we identify feature collapse during initialization and propose quantitative metrics based on feature entropy and minimum eigenvalues. Our analysis reveals the critical roles of token spaces and linear submodules in generalization, systematically validated through multi-scale architectural comparison experiments. Results demonstrate that the proposed surrogate metrics improve correlation rankings with generalization performance by 18%–48%, enabling precise identification of low-compute, high-accuracy ViT architectures. This work provides both theoretical foundations and practical guidance for efficient visual model design.
Multi-view Transformers (e.g., DUSt3R) achieve strong performance in 3D vision but remain poorly understood, hindering interpretability and safe deployment. To address this, we propose the first layer-wise interpretability analysis framework for multi-view Transformers, leveraging residual connection feature probing and geometry-aware visualization to systematically uncover the evolution of internal 3D representations. Our analysis reveals that hidden states progressively encode 3D structure: early layers focus on local correspondences, while deeper layers refine these via geometric reconstruction—enabling end-to-end pose estimation without explicit global pose modeling. This work is the first to elucidate *how* and *why* such models succeed, breaking the “black-box” barrier and providing theoretical foundations for architecture design and reliability validation. Code is publicly available.
Existing light field super-resolution methods struggle to effectively model spatial-angular correlations while preserving 4D ray consistency, primarily due to dimension-decoupled strategies and scanning mechanisms that violate epipolar geometry. To address these limitations, this work proposes SMART, a hybrid network that uniquely integrates angular priors into spatial modeling. SMART synergistically combines a slope-guided Mamba module with an angle-optimized Transformer, complemented by an angle-modulated spatial module and a manifold alignment trajectory module that adheres to epipolar geometry. This architecture enables geometrically consistent cross-dimensional feature learning. Evaluated on five benchmarks, the proposed method achieves state-of-the-art performance, yielding an average PSNR gain of 0.42 dB and significantly suppressing visual artifacts.
This work addresses the challenges of view inconsistency and geometric distortion in single-image 3D reconstruction, which often arise from projection ambiguities during multi-view synthesis. To mitigate these issues, the authors propose a view-adaptive neural rendering framework that employs a shared feature backbone to capture global structure while enabling per-view independent correction of rendering errors. A lightweight self-attention fusion module is introduced to integrate multi-view information and enhance geometric consistency without relying on supervision from diffusion models such as SDS. The method optimizes solely with photometric loss, achieving near state-of-the-art reconstruction fidelity while maintaining computational efficiency and significantly improving view consistency and practical performance.
This study addresses the lack of systematic analysis regarding the geometric evolution of internal representations during Vision Transformer (ViT) training. The authors propose the TGO-II framework, which integrates Centered Kernel Alignment (CKA), Singular Vector Canonical Correlation Analysis (SVCCA), TwoNN intrinsic dimension estimation, and token covariance analysis. Applying this framework to ViT-Small/16 under supervised training, they uncover a tripartite geometric evolution pattern: progressive layer-wise specialization, an initial rise followed by stabilization of intrinsic dimensionality, and the persistent presence of strong token interaction structures. These findings demonstrate that increased representational complexity co-occurs with layer specialization without requiring token decorrelation, thereby challenging the conventional assumption that complexity arises from token independence and highlighting ViT’s capacity to achieve rich representational transformations through sustained strong token interactions.
研究针对多视角视觉Transformer在相机异质性下的相对位置编码问题,提出G-ray方法,通过基于光线角度的旋转相位参数化实现投影不变的位置一致性。
This work addresses the scalability bottleneck in feed-forward novel view synthesis, where Transformer-based architectures suffer from prohibitive computational and memory overheads as the number of input views increases. To this end, we propose the AIMS framework, whose core innovation lies in decoupling the number of observations from the number of processed views via an anchor integration mechanism. Specifically, AIMS employs farthest point sampling to select a fixed set of anchors, combined with spatial clustering and a lightweight learnable integrator to aggregate information from neighboring observations. This design enables the efficient utilization of large-scale multi-view data under a constant global view budget. Experimental results demonstrate that AIMS achieves PSNR values of 29.41 dB on RealEstate10K and 17.73 dB on ScanNet, with an average rendering time of only 7.24 milliseconds per view, yielding a superior quality-efficiency trade-off.