Why Cross-Skeleton Retargeting Is Non-Identifiable: Structural Limits of Generative Motion Models

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the structural unidentifiability problem in cross-skeleton motion retargeting, where generative models struggle to disentangle source motion transfer from canonical pose recovery. We theoretically demonstrate that this ambiguity stems from data sparsity rather than coincidence, revealing the canonical unidentifiability of mappings under standard objectives and the degeneracy of conditional means. By integrating generative motion modeling, unpaired distribution matching, and latent space analysis, we propose Source Instance Fidelity (SIF), a metric quantifying source information preservation that overcomes the blind spots of conventional motion-level evaluation and establishes an observability benchmark. Our findings reveal that existing methods frequently degenerate into source-agnostic baselines on animal datasets, underscoring the urgent need for novel objectives and evaluation frameworks to achieve faithful retargeting.
📝 Abstract
Cross-skeleton motion generation trains generative models to carry action structure and motion intention from one body to another. Yet a target motion that shows the right action has two explanations that the training data cannot tell apart: the model transferred the source clip, or it recovered a typical motion for the requested action. We show that this ambiguity is structural rather than incidental: under standard generative objectives, the source-conditioned retargeting map is non-identifiable in sparse heterogeneous motion domains. Unpaired distribution matching yields gauge non-identifiability: the latent spaces of different skeletons can be transformed relative to one another without changing the training evidence, so different source-conditioned maps fit it equally well. Sparse paired supervision admits the complementary failure mode, \emph{conditional-mean degeneration}: when clips are paired only by action, squared-error training converges to an average target motion that ignores the source clip. To make the missing evidence observable, we introduce Source-Instance Fidelity (SIF), a diagnostic that tests whether outputs differ from one another the way their source clips do, with the target skeleton and action held fixed. Under this diagnostic, methods that succeed at the standard action-level test on animal motion data often sit at the source-blind floor, while the methods that rise above it retain only a partial relational signal. Retargeting therefore needs objectives and evaluations that can identify the source-conditioned map it claims to learn. Project page: https://cross-skeleton-retargeting.netlify.app/.
Problem

Research questions and friction points this paper is trying to address.

cross-skeleton retargeting
non-identifiability
generative motion models
conditional-mean degeneration
distribution matching
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Skeleton Retargeting
Non-Identifiability
Conditional-Mean Degeneration
Source-Instance Fidelity
Generative Motion Models
🔎 Similar Papers
No similar papers found.