pose-conditioned encoding

Designs and implements feature encoders that explicitly incorporate camera or object pose parameters into image or spatial representations, e.g., modules that inject pose vectors into CNN/Transformer layers or embeddings. These pose-conditioned encodings produce pose-aware features used to align information across views, enable pose-conditioned fusion, and support spatial reasoning or correspondence.

pose-conditionedencoding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.44
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Cameras as Relative Positional Encoding

Jul 14, 2025
RL
Ruilong Li
🏛️ UC Berkeley | HKU

Limited 3D perception in multi-view vision tasks stems from insufficient camera geometry modeling. To address this, we propose Projective Positional Encoding (PRoPE), the first method to encode the full camera intrinsic and extrinsic parameters—defining the frustum geometry—as relative positional encodings within Transformers. PRoPE jointly integrates token-level ray-map encoding, attention-level relative pose encoding, and geometrically grounded positional encoding to explicitly model cross-view spatial relationships in self-attention. Crucially, it supports generalization across varying sequence lengths, diverse intrinsic parameter distributions, and out-of-distribution (OOD) camera configurations. Extensive experiments on multi-view image synthesis and stereo depth estimation demonstrate consistent performance gains across model scales; improvements are especially pronounced for long sequences, unseen intrinsics, and OOD scenarios. These results validate the broad efficacy of geometry-aware positional encoding for enhancing multi-view Transformers.

Enhancing multi-view transformers with camera geometry for 3D perceptionImproving novel view synthesis via relative camera conditioning techniquesValidating camera encoding benefits across diverse tasks and model sizes

Efficient Encoder-Free Pose Conditioning and Pose Control for Virtual Try-On

Sep 24, 2025
QL
Qi Li
🏛️ Amazon | University of California, Los Angeles

Virtual Try-On (VTON) suffers from poor pose fidelity, reliance on auxiliary modules (e.g., dedicated encoders or control networks), and difficulty in precise pose guidance. Method: This paper proposes a lightweight, end-to-end pose-fusion approach that eliminates extra modules by performing channel-wise spatial concatenation of pose maps and garment images for direct pose guidance. It introduces pose map representation learning and jointly trains with fine-grained segmentation masks and bounding-box masks to balance pose consistency and garment deformation flexibility. Contribution/Results: Experiments demonstrate significant improvements in pose preservation accuracy and visual realism of synthesized images. The method achieves high-quality try-on results across diverse and complex human poses, establishing a parameter-free, concise, and effective pose-control paradigm for end-to-end VTON.

Balancing pose preservation with flexible pose control in VTONIncorporating pose control into Virtual Try-On without adding parametersSelecting optimal pose representation for accurate product alignment

This work addresses the limited scalability of multi-view Transformers due to performance saturation during training when using camera pose–based positional encoding. The authors identify that coupling rotational and translational components of camera poses within value vectors introduces ambiguity in view representation, hindering model scalability. To resolve this, they propose Decoupled Pose Positional Encoding (DPPE), the first method to explicitly separate rotation and translation in pose encoding while integrating relative positional information. DPPE significantly enhances training stability and generalization, achieving superior novel view synthesis under large-scale settings and demonstrating robustness in extrapolation scenarios—such as increased numbers of input views or changes in scene scale—where prior methods typically degrade.

camera-based positional encodingmulti-view transformersnovel view synthesis

This work addresses the lack of provable correctness guarantees in visual pose estimation for safety-critical applications by proposing a certifiable pose estimation algorithm that integrates physics-driven geometric modeling with learning-based methods. The core innovation lies in introducing a Geometric Generative Model (GGM) combined with neural network reachability analysis to construct a multi-stage, certification-aware pipeline capable of delivering verifiable pose estimates and object detection under conditions ranging from unoccluded to complex occlusion scenarios. The approach is validated on both synthetic and real-world imagery—including event camera data—for planar objects such as traffic signs. Experimental results demonstrate that the estimated poses rigorously satisfy pre-specified certification error bounds, thereby achieving, for the first time, provably robust pose perception suitable for safety-critical deployment.

autonomous systemscertifiable guaranteespose estimation

Positional Prompt Tuning for Efficient 3D Representation Learning

Aug 21, 2024
SZ
Shaochen Zhang
🏛️ Xi'an Jiaotong University | University of Illinois at Urbana-Champaign

Existing point cloud representation learning methods overlook the role of positional encoding (PE), while prevailing parameter-efficient fine-tuning (PEFT) approaches struggle to jointly optimize geometric fidelity and parameter efficiency. To address this, we propose Positional Prompt Tuning (PPT), a lightweight PEFT framework. PPT is the first to formulate high-dimensional positional encodings as learnable prompt embeddings, co-designed with multi-scale patch encodings to construct a feature abstraction module that jointly captures local geometry and global structure. Furthermore, it enables dynamic adapter coupling for PE adaptation within the PEFT paradigm. With only 1.05% trainable parameters, PPT achieves 95.01% classification accuracy on ScanObjectNN OBJ_BG—surpassing state-of-the-art prompt- and adapter-based methods. The implementation is publicly available.

Achieving state-of-the-art accuracy with minimal trainable parametersDeveloping parameter-efficient fine-tuning for point cloud analysisImproving 3D representation learning efficiency through positional encoding

Latest Papers

What's happening recently
View more

This study addresses the limited geometric representational capacity of encoders in existing novel view synthesis methods, which arises from overly powerful decoders and pixel-level objectives. To this end, we propose SNAP, a self-supervised architecture built upon a Transformer encoder-decoder framework. By constraining decoder expressivity to prevent the suppression of geometric structures, and by introducing pose-conditioned local decoding alongside latent-space reconstruction objectives, SNAP optimizes feature learning and endows patch-level representations with emergent viewpoint invariance. Experimental results demonstrate that SNAP achieves performance comparable to specialized supervised models across five tasks, including localization and pose estimation. Furthermore, it significantly outperforms standard 2D representations under camera displacement scenarios, effectively reducing both computational and data requirements.

Decoder ExpressivityGeometric Representation LearningNovel View Synthesis

This study addresses the challenges of spatial orientation control, attribute leakage, and artifact generation in multi-object text-to-image synthesis by proposing a lightweight 2.5D controllable generation framework. The method introduces a standardized OrientLayout dataset alongside a context-aware dual-stream representation and parallel masking architecture. By integrating MM-DiT with sparse spatial angular anchors, the framework achieves precise layout manipulation. Experimental results demonstrate that this approach significantly outperforms state-of-the-art methods in spatial precision, orientation accuracy, and multi-object visual fidelity. Consequently, it effectively resolves persistent issues regarding generation consistency and controllability within complex scenes, offering a robust solution for fine-grained compositional image generation.

2.5D GenerationAttribute LeakageControllable Image Generation

Existing video generation models suffer from poor cross-frame 3D consistency and limited camera controllability due to the absence of explicit 3D structural modeling. To address this, this work proposes RayPE—a novel ray-space positional encoding based on Plücker coordinates—that is, for the first time, additively integrated into the self-attention mechanism of video diffusion Transformers. This integration naturally decomposes attention scores into content, geometry, and their interaction terms. Combined with QK-swapped attention, gating mechanisms, and synergistic normalization using RMSNorm and QKNorm, RayPE achieves substantial improvements in camera controllability, 3D consistency, and overall video quality on mixed multi-view camera data, while introducing less than 0.1% additional parameters.

3D-aware video generationcamera raysPlucker coordinates

This work addresses the limitations of existing novel view synthesis methods, which often lack global and intuitive control over target viewpoints and typically rely on known input camera poses or support only sparse view generation. The authors reformulate novel view synthesis as an image editing task by generating views in the Normalized Object Coordinate Space (NOCS) conditioned on absolute camera poses, eliminating the need for input pose annotations. To enhance generalization, they introduce a text-guided NOCS alignment mechanism that defines a canonical object coordinate system via textual descriptions. Key contributions include the first method enabling global viewpoint control through absolute camera poses, the proposed text-guided NOCS alignment, and a high-quality NOCS dataset. The approach generates consistent, high-fidelity novel views from arbitrary pose-free images across diverse object categories, achieving state-of-the-art performance.

Camera PoseGlobal Pose ControlNormalized Object Coordinate Space

Hot Scholars

XZ

Xiaowei Zhou

Professor of Computer Science, Zhejiang University
Computer VisionComputer Graphics
CS

Chunhua Shen

Zhejiang University
Computer VisionMachine Learning
JM

Jiri Matas

Professor, Czech Technical University
computer visionimage processingpattern recognitionmachine learning
SP

Sida Peng

Zhejiang University
Computer VisionComputer Graphics
SS

Shunsuke Saito

Research Scientist, Meta Codec Avatars Lab
Digital HumansComputer VisionComputer Graphics