🤖 AI Summary
This study addresses a critical limitation in existing developmental robotics simulators, which model infants in isolation and neglect the influence of caregiver interactions on sensorimotor streams. To overcome this, we extend the MIMo platform by introducing a parameterized caregiver model compatible with infant morphology and a source-aware contact logging technique. By replaying natural interactions within the MuJoCo physics engine, our approach generates multimodal tactile and visual data from the infant's first-person perspective. The project successfully produces dense observational data that can be aggregated into touch-rate statistics comparable to manual human coding. Ultimately, this work establishes the first multimodal dataset specifically designed for studying social interaction development, providing a crucial benchmark for embodied cognition research.
📝 Abstract
Early development unfolds in caregiver-infant dyads, where infants'sensorimotor streams are shaped by physical contact and face-to-face interaction. Yet developmental robotics simulators commonly model infants in isolation, limiting the study of caregiver-mediated experience. We present a caregiver-enabled extension of the Multi-Modal Infant Model (MIMo) in MuJoCo that turns MIMo into a controllable platform for replaying dyadic interaction and generating dense infant-perspective observations. The system provides (i) an articulated caregiver model compatible with MIMo morphologies, parameterized from anthropometrics and optionally resized to a recorded caregiver; (ii) a workflow to replay naturalistic caregiver-infant holding and soothing interactions; and (iii) logging and visualizing the infant's first-person multimodal experience (we show touch and vision). We showcase the tool on touch by introducing origin-aware contact logging that disambiguates self-, caregiver-, and environment-generated contact and supports aggregation into touch-rate statistics comparable to manual coding. While we showcase tactile analysis, the platform is intended more broadly as a generator of multimodal dyadic datasets (touch and egocentric vision) for modeling the development of social interaction.