🤖 AI Summary
Current artificial intelligence remains largely confined to digital modalities such as text, vision, and audio, falling short of comprehensively perceiving the rich multisensory world inhabited by humans. This work proposes a ten-year research vision for multisensory intelligence, introducing the first systematic framework that unifies AI with the full spectrum of human senses. Centered on three core directions—sensory perception, modeling, and human-AI collaboration—the framework integrates physiological, tactile, and environmental signals to advance key technologies including cross-modal alignment, unified representation learning, multisensory generation, and transfer learning, thereby overcoming fundamental bottlenecks in heterogeneous modality fusion. The contribution includes a comprehensive research roadmap accompanied by a suite of projects, open-source resources, and demonstration systems developed by the MIT Media Lab team, aiming to catalyze a new paradigm in multimodal human-AI interaction.
📝 Abstract
Our experience of the world is multisensory, spanning a synthesis of language, sight, sound, touch, taste, and smell. Yet, artificial intelligence has primarily advanced in digital modalities like text, vision, and audio. This paper outlines a research vision for multisensory artificial intelligence over the next decade. This new set of technologies can change how humans and AI experience and interact with one another, by connecting AI to the human senses and a rich spectrum of signals from physiological and tactile cues on the body, to physical and social signals in homes, cities, and the environment. We outline how this field must advance through three interrelated themes of sensing, science, and synergy. Firstly, research in sensing should extend how AI captures the world in richer ways beyond the digital medium. Secondly, developing a principled science for quantifying multimodal heterogeneity and interactions, developing unified modeling architectures and representations, and understanding cross-modal transfer. Finally, we present new technical challenges to learn synergy between modalities and between humans and AI, covering multisensory integration, alignment, reasoning, generation, generalization, and experience. Accompanying this vision paper are a series of projects, resources, and demos of latest advances from the Multisensory Intelligence group at the MIT Media Lab, see https://mit-mi.github.io/.