Score
Designs and implements feature encoders that explicitly incorporate camera or object pose parameters into image or spatial representations, e.g., modules that inject pose vectors into CNN/Transformer layers or embeddings. These pose-conditioned encodings produce pose-aware features used to align information across views, enable pose-conditioned fusion, and support spatial reasoning or correspondence.
Limited 3D perception in multi-view vision tasks stems from insufficient camera geometry modeling. To address this, we propose Projective Positional Encoding (PRoPE), the first method to encode the full camera intrinsic and extrinsic parameters—defining the frustum geometry—as relative positional encodings within Transformers. PRoPE jointly integrates token-level ray-map encoding, attention-level relative pose encoding, and geometrically grounded positional encoding to explicitly model cross-view spatial relationships in self-attention. Crucially, it supports generalization across varying sequence lengths, diverse intrinsic parameter distributions, and out-of-distribution (OOD) camera configurations. Extensive experiments on multi-view image synthesis and stereo depth estimation demonstrate consistent performance gains across model scales; improvements are especially pronounced for long sequences, unseen intrinsics, and OOD scenarios. These results validate the broad efficacy of geometry-aware positional encoding for enhancing multi-view Transformers.
Virtual Try-On (VTON) suffers from poor pose fidelity, reliance on auxiliary modules (e.g., dedicated encoders or control networks), and difficulty in precise pose guidance. Method: This paper proposes a lightweight, end-to-end pose-fusion approach that eliminates extra modules by performing channel-wise spatial concatenation of pose maps and garment images for direct pose guidance. It introduces pose map representation learning and jointly trains with fine-grained segmentation masks and bounding-box masks to balance pose consistency and garment deformation flexibility. Contribution/Results: Experiments demonstrate significant improvements in pose preservation accuracy and visual realism of synthesized images. The method achieves high-quality try-on results across diverse and complex human poses, establishing a parameter-free, concise, and effective pose-control paradigm for end-to-end VTON.
This work addresses the limited scalability of multi-view Transformers due to performance saturation during training when using camera pose–based positional encoding. The authors identify that coupling rotational and translational components of camera poses within value vectors introduces ambiguity in view representation, hindering model scalability. To resolve this, they propose Decoupled Pose Positional Encoding (DPPE), the first method to explicitly separate rotation and translation in pose encoding while integrating relative positional information. DPPE significantly enhances training stability and generalization, achieving superior novel view synthesis under large-scale settings and demonstrating robustness in extrapolation scenarios—such as increased numbers of input views or changes in scene scale—where prior methods typically degrade.
This work addresses the lack of provable correctness guarantees in visual pose estimation for safety-critical applications by proposing a certifiable pose estimation algorithm that integrates physics-driven geometric modeling with learning-based methods. The core innovation lies in introducing a Geometric Generative Model (GGM) combined with neural network reachability analysis to construct a multi-stage, certification-aware pipeline capable of delivering verifiable pose estimates and object detection under conditions ranging from unoccluded to complex occlusion scenarios. The approach is validated on both synthetic and real-world imagery—including event camera data—for planar objects such as traffic signs. Experimental results demonstrate that the estimated poses rigorously satisfy pre-specified certification error bounds, thereby achieving, for the first time, provably robust pose perception suitable for safety-critical deployment.
Existing point cloud representation learning methods overlook the role of positional encoding (PE), while prevailing parameter-efficient fine-tuning (PEFT) approaches struggle to jointly optimize geometric fidelity and parameter efficiency. To address this, we propose Positional Prompt Tuning (PPT), a lightweight PEFT framework. PPT is the first to formulate high-dimensional positional encodings as learnable prompt embeddings, co-designed with multi-scale patch encodings to construct a feature abstraction module that jointly captures local geometry and global structure. Furthermore, it enables dynamic adapter coupling for PE adaptation within the PEFT paradigm. With only 1.05% trainable parameters, PPT achieves 95.01% classification accuracy on ScanObjectNN OBJ_BG—surpassing state-of-the-art prompt- and adapter-based methods. The implementation is publicly available.
This study addresses the limited geometric representational capacity of encoders in existing novel view synthesis methods, which arises from overly powerful decoders and pixel-level objectives. To this end, we propose SNAP, a self-supervised architecture built upon a Transformer encoder-decoder framework. By constraining decoder expressivity to prevent the suppression of geometric structures, and by introducing pose-conditioned local decoding alongside latent-space reconstruction objectives, SNAP optimizes feature learning and endows patch-level representations with emergent viewpoint invariance. Experimental results demonstrate that SNAP achieves performance comparable to specialized supervised models across five tasks, including localization and pose estimation. Furthermore, it significantly outperforms standard 2D representations under camera displacement scenarios, effectively reducing both computational and data requirements.
本文提出一种基于几何驱动的数据优化方法,通过主轴对齐解决物体姿态估计中的噪声敏感、对称性混淆问题,且无需修改现有网络架构。
This study addresses the challenges of spatial orientation control, attribute leakage, and artifact generation in multi-object text-to-image synthesis by proposing a lightweight 2.5D controllable generation framework. The method introduces a standardized OrientLayout dataset alongside a context-aware dual-stream representation and parallel masking architecture. By integrating MM-DiT with sparse spatial angular anchors, the framework achieves precise layout manipulation. Experimental results demonstrate that this approach significantly outperforms state-of-the-art methods in spatial precision, orientation accuracy, and multi-object visual fidelity. Consequently, it effectively resolves persistent issues regarding generation consistency and controllability within complex scenes, offering a robust solution for fine-grained compositional image generation.
Existing video generation models suffer from poor cross-frame 3D consistency and limited camera controllability due to the absence of explicit 3D structural modeling. To address this, this work proposes RayPE—a novel ray-space positional encoding based on Plücker coordinates—that is, for the first time, additively integrated into the self-attention mechanism of video diffusion Transformers. This integration naturally decomposes attention scores into content, geometry, and their interaction terms. Combined with QK-swapped attention, gating mechanisms, and synergistic normalization using RMSNorm and QKNorm, RayPE achieves substantial improvements in camera controllability, 3D consistency, and overall video quality on mixed multi-view camera data, while introducing less than 0.1% additional parameters.
This work addresses the limitations of existing novel view synthesis methods, which often lack global and intuitive control over target viewpoints and typically rely on known input camera poses or support only sparse view generation. The authors reformulate novel view synthesis as an image editing task by generating views in the Normalized Object Coordinate Space (NOCS) conditioned on absolute camera poses, eliminating the need for input pose annotations. To enhance generalization, they introduce a text-guided NOCS alignment mechanism that defines a canonical object coordinate system via textual descriptions. Key contributions include the first method enabling global viewpoint control through absolute camera poses, the proposed text-guided NOCS alignment, and a high-quality NOCS dataset. The approach generates consistent, high-fidelity novel views from arbitrary pose-free images across diverse object categories, achieving state-of-the-art performance.