🤖 AI Summary
This study investigates how humans efficiently recognize everyday activities from complex visual inputs and disentangles the critical motion information required for classification versus action reconstruction. By comparing three representation methods—Temporal Movement Primitives (TMP), Legendre polynomial coefficients, and autoencoders—the contributions of static body poses and temporal dynamics are analyzed across videos of 16 daily activities. The findings reveal that static poses alone suffice for high-accuracy activity classification, with nine key joints identified as most discriminative, whereas naturalistic action reconstruction fundamentally relies on temporal dynamics. TMP and Legendre coefficients achieve comparable, top performance in classification; however, only TMP generates perceptually natural motions, highlighting an essential decoupling between the informational demands of action recognition and action generation.
📝 Abstract
Humans recognize movements effortlessly, even from noisy and complex visual input. But what information in the stimulus allows humans to rapidly classify movements? No framework has systematically compared different strategies of movement analysis to address this question. Here, we used videos of 16 daily activities from the MoVi dataset and compared three strategies: Temporal Movement Primitives (TMPs), which decompose movements into weighted sums of temporally smooth basis functions; Legendre polynomial coefficients, which project joint-coordinate trajectories onto an orthogonal polynomial basis; and Autoencoder latent embeddings. Legendre coefficients and TMPs achieved the highest classifier accuracy, followed by autoencoders. We found two discriminative features for movement classification. The most informative is the general posture of the body, the average spatial configuration that distinguishes one activity from another. Additionally, we identified 9 critical joints that are most predictive for movement classification. Interestingly, good classification accuracy did not automatically lead to good movement generation: when we reconstructed movements for each activity, TMPs preserved the temporal dynamics and produced perceptually natural motion, whereas reconstructions from Legendre coefficients retained only the average posture and appeared frozen. These results reveal a dissociation in how movement information is organized: the static configuration of the body suffices to classify what activity is performed, but the temporal dynamics of movement are required to reconstruct how it unfolds. This distinction clarifies which features the visual system may rely upon for rapid action recognition, and suggests that postural features could enable efficient movement screening in clinical applications, while dynamic information remain essential wherever movement generation is the goal.