appearance-pose fusion

Design and implement models and fusion architectures that align and merge visual appearance (e.g., RGB frames or appearance features) with pose/landmark trajectories and motion cues to encode spatio-temporal dynamics. Build encoders, cross-modal aligners, and joint representations that preserve fine-grained pose information (e.g., hands, face) and support downstream prediction and generation tasks.

appearance-posefusion

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

DATA: Domain-And-Time Alignment for High-Quality Feature Fusion in Collaborative Perception

Jul 24, 2025
CT
Chengchang Tian
🏛️ Southeast University | Washington State University

In collaborative perception, hardware heterogeneity induces feature-domain shift, while communication latency causes temporal misalignment—jointly degrading feature quality and accumulating cross-node errors. To address these challenges at the feature-level fusion stage, we propose a systematic alignment framework: (1) a consistency-preserving domain alignment module mitigates inter-device feature distribution discrepancies; (2) a progressive temporal alignment module corrects dynamic timing offsets via multi-scale motion modeling and two-stage compensation; and (3) an observability-constrained discriminator and instance-aware hierarchical aggregation strategy enhance semantic consistency. Evaluated on three benchmark datasets, our method achieves state-of-the-art performance and demonstrates significantly improved robustness under high communication latency and pose estimation errors.

Address domain gaps from hardware diversity and deployment conditionsEnhance semantic feature quality for collaborative perception fusionMitigate temporal misalignment caused by transmission delays

UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework

Dec 03, 2025
YP
Youxin Pang
🏛️ Tsinghua University | Meituan

Existing work primarily focuses on unidirectional cross-modal generation (e.g., text→video or audio→pose), leaving joint modeling of 2D videos and 3D human motion largely unexplored—mainly due to significant structural and distributional heterogeneity between these modalities. To address this gap, we propose UniMo, the first unified autoregressive framework enabling synchronized generation and understanding of 2D videos and 3D human motion. Its core contributions are: (1) mapping both modalities into a shared token sequence via modality-specific embedding layers to mitigate distributional discrepancies; (2) introducing a mixture-of-experts decoder to enhance 3D motion reconstruction fidelity; and (3) designing a VQ-VAE-based 3D motion tokenizer with temporal expansion and vision-token alignment mechanisms. Extensive experiments demonstrate that UniMo achieves state-of-the-art performance across cross-modal synchronized generation, motion capture, and bidirectional understanding tasks.

Addresses structural differences between video and motion dataEnables simultaneous generation and understanding of both modalitiesUnifies 2D video and 3D motion generation in one framework

SMRABooth: Subject and Motion Representation Alignment for Customized Video Generation

Dec 13, 2025
XX
Xuancheng Xu
🏛️ Nanjing University of Posts and Telecommunications | Peng Cheng Laboratory

Existing custom video generation methods struggle to simultaneously preserve subject appearance fidelity and ensure temporal motion consistency, primarily due to the lack of object-level subject-motion disentangled modeling. This paper proposes a subject-motion representation disentanglement and alignment framework for text-to-video generation. We introduce the first object-level representation alignment mechanism, design a sparse spatiotemporal LoRA injection strategy to minimize fine-tuning interference, and develop a collaborative self-supervised subject encoder and optical-flow-based motion encoder. The method integrates self-supervised representation learning, optical-flow-driven motion modeling, efficient LoRA-based fine-tuning, and spatiotemporally sparse adapters. Evaluated on multiple benchmarks, it achieves significant improvements in subject similarity (+12.6%) and motion consistency (+9.8%), enabling fine-grained, disentangled controllable generation with both high visual fidelity and temporal stability.

Ensuring subject appearance similarity in customized video generationMaintaining temporally consistent motion patterns from reference videosReducing interference between subject and motion representations alignment

In pose-guided video generation, simultaneously ensuring appearance consistency (e.g., physique and anatomical proportions) and temporal coherence remains challenging. This paper proposes a training-free enhancement framework to address this issue. Our method introduces (1) structure-motion disentangled modeling, explicitly separating skeletal geometric priors from dynamic motion priors; and (2) pixel-level conditional alignment, jointly leveraging pose-guided geometric calibration and reference-image-driven feature mapping to preserve inter-frame appearance fidelity. This dual-alignment strategy requires no fine-tuning and operates without large-scale annotated data. Experiments demonstrate substantial improvements in physique consistency, proportion stability, and temporal smoothness. Under low-resource settings—without access to task-specific training data or model adaptation—our approach achieves visual quality comparable to supervised training methods. The framework establishes a new paradigm for lightweight, controllable character animation generation, balancing expressiveness, fidelity, and efficiency.

Decouples skeletal and motion priors for precise controlEnsures appearance consistency in character animationsImproves pixel-level alignment for temporal consistency

Existing video customization methods rely on reference images or task-specific temporal priors, which struggle to fully exploit the intrinsic spatiotemporal information in videos, thereby limiting generation flexibility and generalization. This work proposes OmniTransfer, a unified framework that enhances appearance consistency through multi-view inter-frame information and integrates temporal cues for fine-grained temporal control. OmniTransfer introduces three key mechanisms: task-aware positional bias, reference-decoupled causal learning, and task-adaptive multimodal alignment. Notably, it achieves high-quality motion transfer without requiring pose annotations—a first in the field—and unifies support for diverse video transfer tasks. Experiments demonstrate that OmniTransfer outperforms existing approaches in identity and style transfer as well as camera motion and visual effect generation, while matching pose-based models in motion transfer fidelity, enabling highly realistic and flexible video synthesis.

reference imagesspatio-temporal informationtemporal priors

Latest Papers

What's happening recently
View more

This work addresses the challenge in text-driven 3D human motion editing of simultaneously preserving the stylistic characteristics of the source motion while accurately adhering to textual instructions. To this end, the authors propose a dual-axis anchored Transformer architecture that separately extracts features along the joint and temporal dimensions and integrates them through a cross-axis fusion module for joint modeling. A novel auxiliary task—joint-level motion difference prediction—is introduced, leveraging Soft-DTW distance regression to guide the model in identifying critical joints and time steps requiring modification. Integrated within a diffusion-based generative framework and evaluated on the newly curated MotionFix dataset, the method significantly enhances both semantic alignment with input text and fidelity to the original motion structure, achieving state-of-the-art performance.

3D human motionjoint-wise editingmotion preservation

Existing representation alignment methods in diffusion Transformers predominantly rely on point-to-point matching, which struggles to capture the intrinsic spatial structural relationships present in vision foundation models, thereby limiting both training efficiency and generation quality. This work proposes sREPA, a structured representation alignment framework that, for the first time, elevates the alignment granularity from individual points to structural levels. By explicitly modeling geometric consistency of relational structures between feature maps, sREPA effectively preserves the spatial topological structure of pretrained features. Integrated within a diffusion Transformer architecture, sREPA significantly accelerates model convergence, enhances training stability, and achieves superior generation quality compared to existing approaches.

Diffusion TransformersFeature MatchingRelational Geometry

Existing multimodal world models struggle to effectively leverage the rich prior knowledge embedded in foundation models of individual modalities. To address this limitation, this work proposes M²-REPA, a novel approach that introduces, for the first time, a representation alignment mechanism tailored for multimodal video generation. The method decouples modality-specific features from intermediate representations of a diffusion model and aligns them separately with their corresponding foundation models. By integrating modality-decoupling regularization and multimodal alignment losses, M²-REPA enables collaborative optimization of multimodal representations. This approach significantly enhances both visual quality and long-term temporal consistency of generated videos, outperforming current state-of-the-art baselines.

foundation modelsmodality priorsmultimodal world models

This work addresses the challenges of body distortion and facial artifacts commonly encountered in portrait animation when generating long videos or depicting vigorous motions. To mitigate these issues, the authors propose SemanticREPA, a method that integrates human structural and identity (ID) semantic representations as supervisory signals—rather than conditional inputs—within a diffusion model framework. By introducing dedicated structure-alignment and ID-alignment modules, and leveraging video depth estimation alongside face recognition features, the approach utilizes structural priors to guide faithful reconstruction of identity-critical regions. This design effectively enhances both structural stability and identity consistency in the generated outputs. Experimental results demonstrate that SemanticREPA significantly outperforms existing methods under complex motion dynamics and extended temporal generation scenarios.

3D geometric relationshipsfacial distortionhuman image animation

This work addresses the blurring and geometric artifacts in dynamic 3D reconstruction caused by temporal asynchrony in multi-camera systems. The authors propose an asynchronous spatiotemporal reconstruction method supervised by 2D motion trajectories, introducing for the first time an explicit, texture-agnostic 2D trajectory alignment to replace conventional photometric supervision. By jointly optimizing temporal offsets and dynamic 3D representations, the approach effectively mitigates the challenges of missing signals in low-texture regions and the coupling between temporal errors and deformation. Built upon a dynamic Gaussian splatting framework, the method integrates trajectory projection alignment, dynamic masks, and confidence masks to suppress unreliable constraints, significantly enhancing robustness. Experiments demonstrate that the method preserves high-frequency details even under temporal offsets up to 25 frames, achieving a 1.4 dB improvement in PSNR, a 54.0% reduction in temporal offset MAE, and nearly a fourfold increase in synchronization success rate.

asynchronous reconstructiondynamic 3D scenegeometric artifacts

Hot Scholars

MP

Marc Pollefeys

Professor of Computer Science, ETH Zurich, and Director Spatial AI Lab, Microsoft
Computer VisionComputer GraphicsRoboticsMachine Learning
YW

Yi Wan

Pokee AI
reinforcement learning
MR

Martin R. Oswald

University of Amsterdam
3D Computer VisionRepresentation LearningApplied Machine LearningOptimization
JZ

Jiang Zhao

Beihang University
Flight controlAutonomous controlCooperative controlGuidance
CH

Cong Huang

University of Science and Technology of China
Image/Video processing