shape-motion disentanglement

Design and implement representation, encoder–decoder, or generative systems that explicitly separate static shape/appearance (identity, anatomy, spatial structure) from dynamic motion/pose by producing dedicated latent codes or factorized slot representations for each factor. This work includes developing disentanglement objectives, intra-modal and representation distillation techniques, feature extraction and factorization modules, and evaluation/analysis procedures to reduce cross-factor entanglement and enable manipulation, temporally consistent synthesis, restoration, or targeted augmentation of the separated factors.

shape-motiondisentanglement

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.29
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Sculpting Latent Spaces With MMD: Disentanglement With Programmable Priors

Oct 13, 2025
QF
Quentin Fruytier
🏛️ The University of Texas at Austin | InterDigital Communications

This work identifies a fundamental limitation of KL-divergence-based prior regularization in variational autoencoders (VAEs): it fails to reliably enforce posterior aggregation toward a factorized Gaussian prior, resulting in entangled latent representations. To address this, we propose a programmable prior framework grounded in the Maximum Mean Discrepancy (MMD), enabling flexible modeling of complex, semantically aligned priors within the VAE architecture. Furthermore, we introduce an unsupervised Latent Predictability Score (LPS) to quantitatively assess disentanglement. Experiments on CIFAR-10 and Tiny ImageNet demonstrate that our method achieves state-of-the-art mutual information-based disentanglement performance while preserving high-fidelity reconstructions—thereby resolving the inherent reconstruction–disentanglement trade-off prevalent in conventional disentanglement approaches.

Achieving state-of-the-art independence without reconstruction trade-offsAddressing KL divergence's failure in disentangled representation learningIntroducing MMD-based framework to sculpt latent spaces programmatically

Disentanglement with Factor Quantized Variational Autoencoders

Sep 23, 2024
GB
Gulcin Baykal
🏛️ Istanbul Technical University | University of Southern Denmark

Unsupervised disentangled representation learning faces two key challenges: the absence of ground-truth factor annotations and the difficulty in jointly optimizing disentanglement and reconstruction fidelity. To address these, this paper proposes a novel discrete variational autoencoder (VAE) that— for the first time—integrates scalar quantization of latent variables with a globally shared codebook into the VAE framework, while jointly enforcing total correlation (TC) regularization to explicitly constrain statistical independence among latent dimensions. By unifying discrete representation learning with inductive-bias-driven disentanglement optimization, the method achieves state-of-the-art performance on two standard disentanglement metrics—DCI and InfoMEC—outperforming leading unsupervised approaches. Moreover, it significantly improves image reconstruction quality. The implementation is publicly available.

Complex DatasetsData UnderstandingIndependent Component Analysis

Explicitly Disentangled Representations in Object-Centric Learning

Jan 18, 2024
RM
Riccardo Majellaro
🏛️ Leiden University

This work addresses unsupervised object-centric representation learning by explicitly disentangling shape and texture factors of objects in images, thereby enhancing model robustness to structural variations and cross-object generalization. Methodologically, it is the first to impose a predefined dimensional partitioning of the latent space within an object-centric framework, enforcing strict separation between shape and texture subspaces. Building upon Invariant Slot Attention, the approach introduces prior-driven structural constraints and a dual-branch feature disentanglement architecture, enabling controllable texture generation and cross-shape–texture transfer. Evaluated on multiple standard benchmarks, the method achieves significant improvements in disentanglement quality—measured via established metrics—and consistently outperforms existing baselines across diverse downstream tasks, including segmentation, reconstruction, and compositional generalization. These results empirically validate the effectiveness and practicality of explicit latent-space structural design for object-centric learning.

Object Variability RecognitionShape-Texture DisentanglementUnsupervised Machine Learning

This work addresses the challenge of disentangling underlying factors of variation in unsupervised representation learning by introducing Holographic Reduced Representations (HRR) for the first time into this domain. The proposed method models latent variables as vector superpositions of symbol–value pairs and leverages HRR’s unbinding operation as an inductive bias to encourage approximately independent factorized representations. Theoretical analysis derives an upper bound on the information capacity per slot, offering an information-theoretic interpretation of disentanglement. Empirical results demonstrate that the approach outperforms existing baselines in terms of latent traversability and standard disentanglement metrics, while also exhibiting superior robustness to noise and consistently stable reconstruction performance across varying signal-to-noise ratios.

disentanglementfactors of variationholographic reduced representations

Reanimating Images using Neural Representations of Dynamic Stimuli

Jun 04, 2024
JY
Jacob Yeung
🏛️ Carnegie Mellon University

Current computer vision models remain substantially inferior to humans in dynamic motion understanding, especially in realistic, complex scenes. To address this, we propose a brain-inspired video understanding paradigm: (1) decoding fine-grained optical flow from dynamic visual stimuli directly from full-video fMRI responses, enabling the first end-to-end closed-loop reconstruction of high-fidelity videos from whole-brain activity; (2) leveraging a video diffusion model to disentangle static appearance representations from motion generation, and establishing a bidirectional enhancement mechanism between neural motion representations and artificial optical flow predictions; and (3) achieving object-level spatial resolution in motion signal decoding from brain activity. Experiments demonstrate substantial improvements in video-evoked fMRI response prediction accuracy and enable coherent, photorealistic video generation conditioned solely on the initial frame. Our framework provides an interpretable, generalizable neurocomputational foundation for cross-modal dynamic visual modeling.

Decoding visual motion from brain activity using fMRIEnhancing optical flow prediction with brain motion representationReanimating videos from static images via brain-decoded signals

Latest Papers

What's happening recently
View more

This work addresses the limitations of existing disentanglement methods, which rely on strong generative model assumptions and struggle to adapt to modern representation learning frameworks lacking such priors. The authors propose Riemannian Independent Component Analysis (RICA), reframing disentanglement as a second-order geometric property on the data manifold. By introducing the disentanglement tensor and the notion of pointwise disentanglement, RICA eliminates the need for global generative models and independent latent variables. The method integrates Riemannian geometry, the Hessian of the log-likelihood, and Ricci curvature to construct a local disentanglement analysis framework. In controlled experiments with known ground-truth factors, RICA successfully recovers source signals across diverse manifolds and significantly outperforms conventional ICA approaches that depend on coordinate-based representations.

disentanglementgenerative modelsIndependent Component Analysis

This work addresses the challenge of identity leakage and motion distortion in multi-character animation, which arises from entanglement between identity representations and pose dynamics. To resolve this, the paper proposes a diffusion Transformer-based framework capable of animating an arbitrary number of characters. The approach introduces an Instance-Isolated Latent Representation (IILR) and a novel Three-Stage Decoupled Attention (TSDA) mechanism, complemented by an Adaptive Gating Fusion (AGF) module, to achieve precise and spatiotemporally consistent binding between identity and driving poses. This design effectively mitigates identity-pose mismatches and ambiguity in overlapping regions within multi-character scenes, enabling scalable generation of high-fidelity animations with strong identity consistency and controllable motion.

controllabilityidentity entanglementidentity-pose binding

This work addresses the distribution shift arising from morphological differences between humans and robots by proposing a disentangled cross-embodiment video generation framework. The approach decomposes human demonstration videos into two orthogonal latent spaces representing task semantics and embodiment-specific motion. Through dual contrastive learning, mutual information minimization, and orthogonality constraints, the method explicitly disentangles these representations. By integrating parameter-efficient adapters into a frozen video diffusion model, it enables the generation of temporally coherent and morphologically accurate robot execution videos from a single human demonstration. To the best of our knowledge, this is the first method to achieve task–embodiment disentanglement without requiring paired cross-embodiment data, substantially enhancing the scalability and adaptability of robot learning from large-scale in-the-wild human videos.

cross-embodimentdisentangled representationembodiment gap

Hot Scholars

IL

Itai Lang

Postdoctoral Researcher, The University of Chicago
3D Computer Graphics3D Computer VisionDeep Learning
XQ

Xiaojuan Qi

Assistant Professor, The University of Hong Kong
3D VisionDeep learningArtificial IntelligenceMedical Image Analysis
YL

Yebin Liu

Professor, Tsinghua University
Computer GraphicsComputational Photography3D VisionDigital Humans
QH

Qingdong He

Tencent Youtu Lab
Computer visionGenerative AI3D Vision
ZC

Zeyu Cai

Institute of Heavy Ion Physics, Peking University
AI for SciencePlasma PhysicsAI AgentsNumber Theory