🤖 AI Summary
This study addresses the limitations of existing sensor models, specifically the decoupling of perception from action and the restriction to closed action spaces. To overcome these challenges, this work proposes the SLA framework, which leverages natural language as a semantic interface to unify multimodal observations and actions. Methodologically, the authors construct a large-scale SLA benchmark alongside OpenSLA, a unified model that integrates multimodal large language models, hierarchical action modeling, and a polyhedral captioning pipeline. This architecture facilitates hierarchical action prediction, state understanding, and interpretable reasoning. Experimental results demonstrate that the proposed approach surpasses state-of-the-art methods in real-world tasks, including clinical applications. Furthermore, it exhibits robust capabilities in language-guided evidence grounding and zero-shot generalization, highlighting its potential for open-ended embodied perception and interaction.
📝 Abstract
Sensors are useful not only for understanding the world but also for deciding what to do next. Existing sensor models however largely stop at perception: they recognize states or predict outcomes, leaving actions modeled separately through task-specific and often closed label spaces. We introduce Sensor-Language-Action (SLA) modeling, a framework that connects multimodal sensor observations, natural language, and actions within a unified model. SLA uses language as a semantic interface between sensing and acting, allowing heterogeneous actions to be represented, predicted, and explained while remaining grounded in the underlying sensor evidence. We build a large-scale SLA benchmark consisting of datasets that span more than 116,000 individuals, 79 sensor modalities, and 60 action groups, together with a multi-faceted captioning pipeline that aligns user context, sensor dynamics, and action evidence. Building on this framework, we present OpenSLA, a unified SLA model for hierarchical action prediction, state understanding, and action explanation. Extensive experiments on real-world tasks in clinical prediction, operating rooms, and metabolic health verify its superior performance over the state-of-the-art. OpenSLA also demonstrates intriguing capabilities including language-guided evidence grounding and zero-shot generalization to unseen actions and cohorts.