Score
Design and implement representation, encoder–decoder, or generative systems that explicitly separate static shape/appearance (identity, anatomy, spatial structure) from dynamic motion/pose by producing dedicated latent codes or factorized slot representations for each factor. This work includes developing disentanglement objectives, intra-modal and representation distillation techniques, feature extraction and factorization modules, and evaluation/analysis procedures to reduce cross-factor entanglement and enable manipulation, temporally consistent synthesis, restoration, or targeted augmentation of the separated factors.
This work identifies a fundamental limitation of KL-divergence-based prior regularization in variational autoencoders (VAEs): it fails to reliably enforce posterior aggregation toward a factorized Gaussian prior, resulting in entangled latent representations. To address this, we propose a programmable prior framework grounded in the Maximum Mean Discrepancy (MMD), enabling flexible modeling of complex, semantically aligned priors within the VAE architecture. Furthermore, we introduce an unsupervised Latent Predictability Score (LPS) to quantitatively assess disentanglement. Experiments on CIFAR-10 and Tiny ImageNet demonstrate that our method achieves state-of-the-art mutual information-based disentanglement performance while preserving high-fidelity reconstructions—thereby resolving the inherent reconstruction–disentanglement trade-off prevalent in conventional disentanglement approaches.
Unsupervised disentangled representation learning faces two key challenges: the absence of ground-truth factor annotations and the difficulty in jointly optimizing disentanglement and reconstruction fidelity. To address these, this paper proposes a novel discrete variational autoencoder (VAE) that— for the first time—integrates scalar quantization of latent variables with a globally shared codebook into the VAE framework, while jointly enforcing total correlation (TC) regularization to explicitly constrain statistical independence among latent dimensions. By unifying discrete representation learning with inductive-bias-driven disentanglement optimization, the method achieves state-of-the-art performance on two standard disentanglement metrics—DCI and InfoMEC—outperforming leading unsupervised approaches. Moreover, it significantly improves image reconstruction quality. The implementation is publicly available.
This work addresses unsupervised object-centric representation learning by explicitly disentangling shape and texture factors of objects in images, thereby enhancing model robustness to structural variations and cross-object generalization. Methodologically, it is the first to impose a predefined dimensional partitioning of the latent space within an object-centric framework, enforcing strict separation between shape and texture subspaces. Building upon Invariant Slot Attention, the approach introduces prior-driven structural constraints and a dual-branch feature disentanglement architecture, enabling controllable texture generation and cross-shape–texture transfer. Evaluated on multiple standard benchmarks, the method achieves significant improvements in disentanglement quality—measured via established metrics—and consistently outperforms existing baselines across diverse downstream tasks, including segmentation, reconstruction, and compositional generalization. These results empirically validate the effectiveness and practicality of explicit latent-space structural design for object-centric learning.
This work addresses the challenge of disentangling underlying factors of variation in unsupervised representation learning by introducing Holographic Reduced Representations (HRR) for the first time into this domain. The proposed method models latent variables as vector superpositions of symbol–value pairs and leverages HRR’s unbinding operation as an inductive bias to encourage approximately independent factorized representations. Theoretical analysis derives an upper bound on the information capacity per slot, offering an information-theoretic interpretation of disentanglement. Empirical results demonstrate that the approach outperforms existing baselines in terms of latent traversability and standard disentanglement metrics, while also exhibiting superior robustness to noise and consistently stable reconstruction performance across varying signal-to-noise ratios.
Current computer vision models remain substantially inferior to humans in dynamic motion understanding, especially in realistic, complex scenes. To address this, we propose a brain-inspired video understanding paradigm: (1) decoding fine-grained optical flow from dynamic visual stimuli directly from full-video fMRI responses, enabling the first end-to-end closed-loop reconstruction of high-fidelity videos from whole-brain activity; (2) leveraging a video diffusion model to disentangle static appearance representations from motion generation, and establishing a bidirectional enhancement mechanism between neural motion representations and artificial optical flow predictions; and (3) achieving object-level spatial resolution in motion signal decoding from brain activity. Experiments demonstrate substantial improvements in video-evoked fMRI response prediction accuracy and enable coherent, photorealistic video generation conditioned solely on the initial frame. Our framework provides an interpretable, generalizable neurocomputational foundation for cross-modal dynamic visual modeling.
This work addresses the limitations of existing disentanglement methods, which rely on strong generative model assumptions and struggle to adapt to modern representation learning frameworks lacking such priors. The authors propose Riemannian Independent Component Analysis (RICA), reframing disentanglement as a second-order geometric property on the data manifold. By introducing the disentanglement tensor and the notion of pointwise disentanglement, RICA eliminates the need for global generative models and independent latent variables. The method integrates Riemannian geometry, the Hessian of the log-likelihood, and Ricci curvature to construct a local disentanglement analysis framework. In controlled experiments with known ground-truth factors, RICA successfully recovers source signals across diverse manifolds and significantly outperforms conventional ICA approaches that depend on coordinate-based representations.
This work addresses the challenge of identity leakage and motion distortion in multi-character animation, which arises from entanglement between identity representations and pose dynamics. To resolve this, the paper proposes a diffusion Transformer-based framework capable of animating an arbitrary number of characters. The approach introduces an Instance-Isolated Latent Representation (IILR) and a novel Three-Stage Decoupled Attention (TSDA) mechanism, complemented by an Adaptive Gating Fusion (AGF) module, to achieve precise and spatiotemporally consistent binding between identity and driving poses. This design effectively mitigates identity-pose mismatches and ambiguity in overlapping regions within multi-character scenes, enabling scalable generation of high-fidelity animations with strong identity consistency and controllable motion.
This work addresses the distribution shift arising from morphological differences between humans and robots by proposing a disentangled cross-embodiment video generation framework. The approach decomposes human demonstration videos into two orthogonal latent spaces representing task semantics and embodiment-specific motion. Through dual contrastive learning, mutual information minimization, and orthogonality constraints, the method explicitly disentangles these representations. By integrating parameter-efficient adapters into a frozen video diffusion model, it enables the generation of temporally coherent and morphologically accurate robot execution videos from a single human demonstration. To the best of our knowledge, this is the first method to achieve task–embodiment disentanglement without requiring paired cross-embodiment data, substantially enhancing the scalability and adaptability of robot learning from large-scale in-the-wild human videos.
研究解决了使用欧氏坐标表示生成因素时的几何不匹配问题,提出Factor-Space Topographic Map (FactoMap)方法来学习与因素空间结构相匹配的表示,从而实现因素解耦。
为解决视频生成中缺乏显式语义结构问题,提出SlotDiT,通过将场景分解为基于对象的槽来指导扩散变换器,在机器人应用中提高任务完成率。