🤖 AI Summary
This study addresses the limitations of existing social motion models regarding implicit interaction states and insufficient cross-group transferability by proposing BRAID, a hierarchical sequential latent variable model. This method formulates social motion generation as a meta-transfer learning problem, introducing a novel dual-level representation architecture that explicitly disentangles group and individual dynamics. By leveraging shared interaction priors, context adaptation mechanisms, and a unified SMPL representation framework, it achieves coherent generation under fully sparse observations. Experimental results demonstrate that the proposed model significantly improves reconstruction accuracy, realism, and interpersonal coordination in tasks such as prediction and tracking, thereby validating the advantages of the disentangled structure within the hierarchical latent space.
📝 Abstract
Human social behaviour is not a collection of independent motions, but a jointly organised process in which group dynamics and individual variation continuously shape one another. Yet existing social motion models often prioritise plausible trajectories while leaving interaction state implicit, limiting their ability to transfer across groups, tasks, and partial-observation regimes. To address this gap, we introduce Bilevel Representations for Agent Interaction Dynamics (BRAID), a hierarchical sequential latent-variable model for generative multi-person interaction. BRAID explicitly formulates social motion generation as a meta-transfer learning problem: shared interaction priors are learned across datasets and adapted through arbitrary context sets of observed people and joints. The model represents each scene through a group-level latent state that captures shared interaction dynamics and person-level latent states that capture individual behaviour conditioned on the evolving group context. This modelling choice enables coherent generation under full, sparse, or partial observations while exposing compact social-state vectors that can serve as an interface for downstream embodied-agent systems. We evaluate BRAID under a unified SMPL-based representation on social forecasting, tracking and in-filling, and response generation, using metrics that assess not only reconstruction accuracy but also realism, diversity, temporal alignment, and interpersonal coordination. We further analyse the hierarchical latent space, showing that it captures separable group- and individual-level structure.