Score
Designs and implements pose encoding schemes that represent camera or object pose as disentangled embeddings—e.g., separate rotation and translation channels, camera-aware or camera-based positional encodings, or dedicated pose positional encoding vectors—by allocating distinct embedding dimensions and encoding functions. Builds the accompanying architecture and training practices to preserve independent identification of pose components, stabilize long-term training, and improve generalization to novel viewpoints.
This work addresses the limited scalability of multi-view Transformers due to performance saturation during training when using camera pose–based positional encoding. The authors identify that coupling rotational and translational components of camera poses within value vectors introduces ambiguity in view representation, hindering model scalability. To resolve this, they propose Decoupled Pose Positional Encoding (DPPE), the first method to explicitly separate rotation and translation in pose encoding while integrating relative positional information. DPPE significantly enhances training stability and generalization, achieving superior novel view synthesis under large-scale settings and demonstrating robustness in extrapolation scenarios—such as increased numbers of input views or changes in scene scale—where prior methods typically degrade.
Limited 3D perception in multi-view vision tasks stems from insufficient camera geometry modeling. To address this, we propose Projective Positional Encoding (PRoPE), the first method to encode the full camera intrinsic and extrinsic parameters—defining the frustum geometry—as relative positional encodings within Transformers. PRoPE jointly integrates token-level ray-map encoding, attention-level relative pose encoding, and geometrically grounded positional encoding to explicitly model cross-view spatial relationships in self-attention. Crucially, it supports generalization across varying sequence lengths, diverse intrinsic parameter distributions, and out-of-distribution (OOD) camera configurations. Extensive experiments on multi-view image synthesis and stereo depth estimation demonstrate consistent performance gains across model scales; improvements are especially pronounced for long sequences, unseen intrinsics, and OOD scenarios. These results validate the broad efficacy of geometry-aware positional encoding for enhancing multi-view Transformers.
In controllable video generation, camera motion and object motion are difficult to disentangle due to their shared inverse-depth scaling behavior in optical flow, leading to mutual interference between control signals. This work proposes a geometry–semantics orthogonal attention mechanism that explicitly decouples these factors at the architectural level. Specifically, camera motion is modeled in the geometric branch via norm-preserving rotations based on RoPE phase encoding, while object motion is handled in the semantic branch through gated value injection. A lightweight orthogonality regularization term is introduced to enforce subspace orthogonality between the two response pathways. This constructive design ensures disentanglement by architecture rather than relying on training stochasticity. Experiments demonstrate that the method reduces control crosstalk by over 2.4× while preserving generation fidelity, achieving state-of-the-art accuracy in both camera and object motion control, and generalizing effectively across different backbone architectures.
Existing camera encoding methods rely heavily on the pinhole camera assumption, limiting their generalizability to real-world cameras with complex intrinsic parameters and lens distortions. Method: We propose Unified Camera Positional Encoding (UCPE), the first method to jointly model full geometric information—including 6-DoF pose, intrinsics, radial/tangential distortion, pitch, and roll—via relative ray encoding for light-path characterization and absolute direction encoding for global orientation. UCPE introduces <1% additional trainable parameters and is integrated into a pre-trained video diffusion Transformer with a lightweight spatial attention adapter, trained on a large-scale, in-house dataset covering diverse camera motions and lens types. Contribution/Results: Our approach achieves state-of-the-art performance in camera-controllable video generation, significantly improving visual fidelity and geometric consistency. It demonstrates strong generalization potential across multi-view, video, and 3D tasks, establishing UCPE as a versatile, geometry-aware camera representation.
Current text-to-video (T2V) and image-to-video (I2V) models lack explicit, content-decoupled motion representations, severely limiting motion transfer and editing capabilities. To address this, we propose a self-supervised, alignment-free method for learning abstract motion representations: by reconstructing targets in image space and incorporating a lightweight adapter, motion is fully disentangled from static factors—including appearance, identity, and pose. Our representation enables open-world motion transfer across semantically disparate categories without requiring pixel- or instance-level correspondences, and can be seamlessly integrated—plug-and-play—into arbitrary video generators. In zero-shot action classification, our method significantly outperforms state-of-the-art representation models such as V-JEPA, achieving new SOTA results on Something-Something v2 and Jester. Crucially, it preserves both motion fidelity and text-video alignment.
本文提出一种基于几何驱动的数据优化方法,通过主轴对齐解决物体姿态估计中的噪声敏感、对称性混淆问题,且无需修改现有网络架构。
This study addresses the limited geometric representational capacity of encoders in existing novel view synthesis methods, which arises from overly powerful decoders and pixel-level objectives. To this end, we propose SNAP, a self-supervised architecture built upon a Transformer encoder-decoder framework. By constraining decoder expressivity to prevent the suppression of geometric structures, and by introducing pose-conditioned local decoding alongside latent-space reconstruction objectives, SNAP optimizes feature learning and endows patch-level representations with emergent viewpoint invariance. Experimental results demonstrate that SNAP achieves performance comparable to specialized supervised models across five tasks, including localization and pose estimation. Furthermore, it significantly outperforms standard 2D representations under camera displacement scenarios, effectively reducing both computational and data requirements.
This study addresses the feature mismatch between pre-trained 3D encoders designed for global scenes and the local observations encountered by embodied agents, proposing a pioneering label-free adaptation paradigm. The method freezes a global 3D encoder and trains only a lightweight adapter module using paired geometric information. Through point-feature alignment and relational distillation, it maps local-view features into the global semantic space, enabling global guidance during training while directly processing local observations at inference. Experiments demonstrate that this approach outperforms supervised parameter-efficient fine-tuning on the Sonata and Concerto benchmarks. Furthermore, its zero-shot cross-dataset transfer performance significantly surpasses that of fully fine-tuned models.
Existing video generation models suffer from poor cross-frame 3D consistency and limited camera controllability due to the absence of explicit 3D structural modeling. To address this, this work proposes RayPE—a novel ray-space positional encoding based on Plücker coordinates—that is, for the first time, additively integrated into the self-attention mechanism of video diffusion Transformers. This integration naturally decomposes attention scores into content, geometry, and their interaction terms. Combined with QK-swapped attention, gating mechanisms, and synergistic normalization using RMSNorm and QKNorm, RayPE achieves substantial improvements in camera controllability, 3D consistency, and overall video quality on mixed multi-view camera data, while introducing less than 0.1% additional parameters.
This work addresses the opacity of semantic representations in existing vision-language models and the limitations of conventional sparse autoencoders, which rely on overcomplete expansions that distort geometric structure and introduce redundancy. The authors propose CEDAR, a novel method that—without altering the original embedding dimension—transforms pretrained embeddings into axis-aligned, disentangled representations through an invertible linear transformation, an adaptive rotation mechanism, and a top-k sparsity bottleneck. By avoiding overcompleteness, CEDAR effectively uncovers the compositional structure of embeddings, enabling high-quality concept alignment and natural language decoding in models such as CLIP and BLIP. Experiments demonstrate that CEDAR achieves an excellent trade-off between reconstruction fidelity and sparsity, yielding interpretations that are not only more human-readable but also highly consistent with human cognition.