decouple pose encoding

Designs and implements pose encoding schemes that represent camera or object pose as disentangled embeddings—e.g., separate rotation and translation channels, camera-aware or camera-based positional encodings, or dedicated pose positional encoding vectors—by allocating distinct embedding dimensions and encoding functions. Builds the accompanying architecture and training practices to preserve independent identification of pose components, stabilize long-term training, and improve generalization to novel viewpoints.

decoupleposeencoding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limited scalability of multi-view Transformers due to performance saturation during training when using camera pose–based positional encoding. The authors identify that coupling rotational and translational components of camera poses within value vectors introduces ambiguity in view representation, hindering model scalability. To resolve this, they propose Decoupled Pose Positional Encoding (DPPE), the first method to explicitly separate rotation and translation in pose encoding while integrating relative positional information. DPPE significantly enhances training stability and generalization, achieving superior novel view synthesis under large-scale settings and demonstrating robustness in extrapolation scenarios—such as increased numbers of input views or changes in scene scale—where prior methods typically degrade.

camera-based positional encodingmulti-view transformersnovel view synthesis

Cameras as Relative Positional Encoding

Jul 14, 2025
RL
Ruilong Li
🏛️ UC Berkeley | HKU

Limited 3D perception in multi-view vision tasks stems from insufficient camera geometry modeling. To address this, we propose Projective Positional Encoding (PRoPE), the first method to encode the full camera intrinsic and extrinsic parameters—defining the frustum geometry—as relative positional encodings within Transformers. PRoPE jointly integrates token-level ray-map encoding, attention-level relative pose encoding, and geometrically grounded positional encoding to explicitly model cross-view spatial relationships in self-attention. Crucially, it supports generalization across varying sequence lengths, diverse intrinsic parameter distributions, and out-of-distribution (OOD) camera configurations. Extensive experiments on multi-view image synthesis and stereo depth estimation demonstrate consistent performance gains across model scales; improvements are especially pronounced for long sequences, unseen intrinsics, and OOD scenarios. These results validate the broad efficacy of geometry-aware positional encoding for enhancing multi-view Transformers.

Enhancing multi-view transformers with camera geometry for 3D perceptionImproving novel view synthesis via relative camera conditioning techniquesValidating camera encoding benefits across diverse tasks and model sizes

In controllable video generation, camera motion and object motion are difficult to disentangle due to their shared inverse-depth scaling behavior in optical flow, leading to mutual interference between control signals. This work proposes a geometry–semantics orthogonal attention mechanism that explicitly decouples these factors at the architectural level. Specifically, camera motion is modeled in the geometric branch via norm-preserving rotations based on RoPE phase encoding, while object motion is handled in the semantic branch through gated value injection. A lightweight orthogonality regularization term is introduced to enforce subspace orthogonality between the two response pathways. This constructive design ensures disentanglement by architecture rather than relying on training stochasticity. Experiments demonstrate that the method reduces control crosstalk by over 2.4× while preserving generation fidelity, achieving state-of-the-art accuracy in both camera and object motion control, and generalizing effectively across different backbone architectures.

camera motioncontrollable video generationmotion disentanglement

Unified Camera Positional Encoding for Controlled Video Generation

Dec 08, 2025
CZ
Cheng Zhang
🏛️ Monash University | V AST

Existing camera encoding methods rely heavily on the pinhole camera assumption, limiting their generalizability to real-world cameras with complex intrinsic parameters and lens distortions. Method: We propose Unified Camera Positional Encoding (UCPE), the first method to jointly model full geometric information—including 6-DoF pose, intrinsics, radial/tangential distortion, pitch, and roll—via relative ray encoding for light-path characterization and absolute direction encoding for global orientation. UCPE introduces <1% additional trainable parameters and is integrated into a pre-trained video diffusion Transformer with a lightweight spatial attention adapter, trained on a large-scale, in-house dataset covering diverse camera motions and lens types. Contribution/Results: Our approach achieves state-of-the-art performance in camera-controllable video generation, significantly improving visual fidelity and geometric consistency. It demonstrates strong generalization potential across multi-view, video, and 3D tasks, establishing UCPE as a versatile, geometry-aware camera representation.

Enable full control over camera orientation in video generationIntegrate camera representation into Transformers with minimal parametersUnify camera encoding for diverse intrinsics and distortions

DisMo: Disentangled Motion Representations for Open-World Motion Transfer

Nov 28, 2025
TR
Thomas Ressler-Antal
🏛️ CompVis | LMU Munich | Munich Center for Machine Learning

Current text-to-video (T2V) and image-to-video (I2V) models lack explicit, content-decoupled motion representations, severely limiting motion transfer and editing capabilities. To address this, we propose a self-supervised, alignment-free method for learning abstract motion representations: by reconstructing targets in image space and incorporating a lightweight adapter, motion is fully disentangled from static factors—including appearance, identity, and pose. Our representation enables open-world motion transfer across semantically disparate categories without requiring pixel- or instance-level correspondences, and can be seamlessly integrated—plug-and-play—into arbitrary video generators. In zero-shot action classification, our method significantly outperforms state-of-the-art representation models such as V-JEPA, achieving new SOTA results on Something-Something v2 and Jester. Crucially, it preserves both motion fidelity and text-video alignment.

Enabling motion transfer across semantically unrelated entities without correspondencesOvercoming trade-offs between motion fidelity and prompt adherence in video generationSeparating motion representation from content in video generation models

Latest Papers

What's happening recently
View more

This study addresses the limited geometric representational capacity of encoders in existing novel view synthesis methods, which arises from overly powerful decoders and pixel-level objectives. To this end, we propose SNAP, a self-supervised architecture built upon a Transformer encoder-decoder framework. By constraining decoder expressivity to prevent the suppression of geometric structures, and by introducing pose-conditioned local decoding alongside latent-space reconstruction objectives, SNAP optimizes feature learning and endows patch-level representations with emergent viewpoint invariance. Experimental results demonstrate that SNAP achieves performance comparable to specialized supervised models across five tasks, including localization and pose estimation. Furthermore, it significantly outperforms standard 2D representations under camera displacement scenarios, effectively reducing both computational and data requirements.

Decoder ExpressivityGeometric Representation LearningNovel View Synthesis

This study addresses the feature mismatch between pre-trained 3D encoders designed for global scenes and the local observations encountered by embodied agents, proposing a pioneering label-free adaptation paradigm. The method freezes a global 3D encoder and trains only a lightweight adapter module using paired geometric information. Through point-feature alignment and relational distillation, it maps local-view features into the global semantic space, enabling global guidance during training while directly processing local observations at inference. Experiments demonstrate that this approach outperforms supervised parameter-efficient fine-tuning on the Sonata and Concerto benchmarks. Furthermore, its zero-shot cross-dataset transfer performance significantly surpasses that of fully fine-tuned models.

3D representation mismatchcoordinate frame discrepancyembodied AI

Existing video generation models suffer from poor cross-frame 3D consistency and limited camera controllability due to the absence of explicit 3D structural modeling. To address this, this work proposes RayPE—a novel ray-space positional encoding based on Plücker coordinates—that is, for the first time, additively integrated into the self-attention mechanism of video diffusion Transformers. This integration naturally decomposes attention scores into content, geometry, and their interaction terms. Combined with QK-swapped attention, gating mechanisms, and synergistic normalization using RMSNorm and QKNorm, RayPE achieves substantial improvements in camera controllability, 3D consistency, and overall video quality on mixed multi-view camera data, while introducing less than 0.1% additional parameters.

3D-aware video generationcamera raysPlucker coordinates

This work addresses the opacity of semantic representations in existing vision-language models and the limitations of conventional sparse autoencoders, which rely on overcomplete expansions that distort geometric structure and introduce redundancy. The authors propose CEDAR, a novel method that—without altering the original embedding dimension—transforms pretrained embeddings into axis-aligned, disentangled representations through an invertible linear transformation, an adaptive rotation mechanism, and a top-k sparsity bottleneck. By avoiding overcompleteness, CEDAR effectively uncovers the compositional structure of embeddings, enabling high-quality concept alignment and natural language decoding in models such as CLIP and BLIP. Experiments demonstrate that CEDAR achieves an excellent trade-off between reconstruction fidelity and sparsity, yielding interpretations that are not only more human-readable but also highly consistent with human cognition.

embedding interpretabilitymultimodal embeddingssemantic disentanglement

Hot Scholars

YS

Yujun Shen

Ant Group
Generative ModelingComputer VisionDeep Learning
DT

Dacheng Tao

Nanyang Technological University
artificial intelligencemachine learningcomputer visionimage processing
DC

Daniel Cremers

Technical University of Munich
Computer VisionMachine LearningOptimizationRobotics
FT

Federico Tombari

Google, TU Munich
Computer VisionMachine Learning3D Perception
JY

Jingyi Yu

Professor, ShanghaiTech University
Computer VisionComputer Graphics