Score
Designs, builds, and evaluates algorithms and systems that acquire and process sensory data to detect, segment, track, and infer properties of objects, agents, and environments, producing structured representations such as labels, poses, maps, and state estimates. Work covers modalities like images, video, audio, and range sensors and includes feature extraction, sensor fusion, and uncertainty estimation.
Current artificial intelligence remains largely confined to digital modalities such as text, vision, and audio, falling short of comprehensively perceiving the rich multisensory world inhabited by humans. This work proposes a ten-year research vision for multisensory intelligence, introducing the first systematic framework that unifies AI with the full spectrum of human senses. Centered on three core directions—sensory perception, modeling, and human-AI collaboration—the framework integrates physiological, tactile, and environmental signals to advance key technologies including cross-modal alignment, unified representation learning, multisensory generation, and transfer learning, thereby overcoming fundamental bottlenecks in heterogeneous modality fusion. The contribution includes a comprehensive research roadmap accompanied by a suite of projects, open-source resources, and demonstration systems developed by the MIT Media Lab team, aiming to catalyze a new paradigm in multimodal human-AI interaction.
Current machine perception ecosystems are predominantly vision- and audition-centric, lacking high-sensitivity, high-specificity molecular-scale chemical sensing capabilities. To address this gap, this project introduces the “Global Chemical Sensing Infrastructure” paradigm: it stabilizes mammalian olfactory receptors and integrates them into biophotonic/electronic transduction platforms, augmented by embedded multimodal AI and a distributed sensor network—enabling, for the first time, machine olfaction approaching single-molecule resolution. The system achieves detection accuracy and response latency comparable to professional detection dogs, supporting non-invasive, real-time odor monitoring. This advance enables intelligent sensing upgrades across clinical early diagnosis, industrial process control, precision agriculture, and public safety—establishing a novel molecular perception layer spanning health, environment, and security domains, while catalyzing interdisciplinary applications and emerging markets.
To address the challenges of real-time visual perception and decision-making in complex dynamic environments, this work proposes an active visual perception framework that transcends traditional passive vision paradigms by enabling sensor actuation and attention-driven control to close the perception–action loop. Methodologically, it integrates computer vision, deep reinforcement learning, multimodal sensor fusion, and a lightweight real-time decision module into an end-to-end trainable active perception system. Key contributions include: (1) an online attention-guidance policy conditioned on environmental feedback; (2) a tightly coupled spatiotemporal alignment mechanism for heterogeneous multi-sensor data; and (3) a low-latency closed-loop control architecture. Experiments on robotic navigation, autonomous driving simulation, and interactive tasks demonstrate a 37% improvement in perception efficiency and a 52% reduction in decision latency, significantly enhancing adaptability and robustness in dynamic scenarios.
High-level autonomous driving (HAD) systems suffer from poor cross-hardware generalization due to multi-modal perception models’ strong dependence on specific sensor hardware configurations. Method: This paper proposes the first sensor data abstraction framework tailored for HAD, systematically defining and implementing unified abstraction interfaces for cameras, LiDAR, and millimeter-wave radar. Integrating signal processing, geometric modeling, and representation learning, the framework enables hardware-agnostic representations for both uni-modal and multi-modal fusion. Contribution/Results: Evaluated on diverse real-world datasets, this work identifies— for the first time—the core challenges and technical pathways for abstracting all three sensor modalities. It establishes a theoretical foundation and architectural blueprint for building scalable, generalizable perception models that transcend hardware-specific constraints, thereby advancing robust, deployment-ready HAD systems.
Existing visual SLAM methods exhibit severe generalization deficits across diverse applications (e.g., XR, IoT, autonomous driving, UAVs, human pose tracking) and heterogeneous environments (indoor/outdoor, static/dynamic scenes, varying motion patterns), stemming from deep coupling among algorithm design, environmental characteristics, and platform motion dynamics. Method: We propose the first three-dimensional challenge taxonomy—“algorithm–environment–motion”—and systematically evaluate state-of-the-art methods (ORB-SLAM2/3, VINS-Fusion) on multi-source benchmarks (TUM, EuRoC, ARKitScenes, UAV-Human), quantifying performance via absolute trajectory error (ATE), relative pose error (RPE), and tracking loss. Contribution/Results: No method achieves robust cross-domain or intra-domain heterogeneous generalization. To address this, we introduce a principled co-optimization pathway comprising input representation disentanglement, intermediate information reuse, and output dynamic validation—establishing a reproducible benchmark and foundational design principles for universal visual localization.
This work addresses the lack of effective auditing mechanisms in current AI-based motion capture systems, which hinders verification of whether inferred skeletal poses conform to authentic human behavior. To tackle this challenge, the authors propose a novel auditing framework that integrates contextual reasoning with biomechanical symmetry principles. By embedding practical application contexts into the evaluation process and combining motion capture outputs with empirically verifiable real-world measurements—even in the absence of ground truth or amid contested annotations—the method enables rigorous empirical auditing of system behavior. The approach successfully uncovers implicit assumptions and biases in how ground truth is defined within existing systems, thereby offering both theoretical foundations and practical pathways for trustworthy assessment of motion capture technologies.
Traditional robotic perception is constrained by fixed or onboard sensors, limiting the ability to flexibly acquire task-optimal viewpoints. This work proposes SensorPerch, which decouples sensing from the robot body and environmental infrastructure by introducing autonomous, deployable, and retrievable sensor units as independent physical entities. Built upon a lightweight, wireless, reconfigurable sensor platform and a task-driven viewpoint selection framework, SensorPerch enables stable attachment to diverse surfaces and on-demand sensor placement. Experimental results demonstrate its effectiveness in both object-coupled and policy-coupled tasks, achieving persistent long-range state monitoring and policy success rates comparable to those obtained with prior knowledge of the optimal viewpoint.
This work addresses the challenge of event localization in communication-denied environments, where path-integral sensors provide only binary path observations, thereby hindering precise event detection and limiting information fusion and path planning. The paper proposes a Bayesian network–based belief map updating method that, for the first time, enables principled Bayesian inference over path observations by explicitly modeling false alarms and missed detections. By integrating Shannon information theory, the approach plans trajectories that maximize information gain. In contrast to existing methods relying on posterior mean approximations, the proposed technique significantly accelerates belief map convergence and substantially improves both accuracy and efficiency in static hazard detection, demonstrating consistent advantages in both single-robot and multi-robot scenarios.
This work addresses the challenges of high early-stage uncertainty in manufacturing monitoring system development—leading to redundant modeling and substantial training costs—and the limited transferability of filtering pipelines in cross-domain image segmentation tasks. To tackle these issues, the authors propose a problem-centric design paradigm that constructs an abstract system model to continuously accumulate and retrieve historical segmentation tasks along with their associated filtering pipelines, enabling solution reuse and incremental optimization. The approach integrates similarity-based problem retrieval, abstract modeling, pipeline reuse, and a retrieval-augmented evolutionary learning mechanism. Experimental results demonstrate that the method significantly reduces training costs and late-stage revision risks, provides the first systematic validation of filtering pipeline transferability across similar segmentation tasks, and achieves a favorable balance among complexity, technical requirements, and reliability under lightweight model constraints.
This work addresses the labor-intensive and expertise-dependent nature of computational imaging system design by proposing a method that automatically generates verifiable forward models from natural language instructions. Leveraging a formal specification language (spec.md) and a multi-agent architecture comprising Plan, Judge, and Execute modules, the approach combines a finite primitive basis to translate single-sentence descriptions into imaging systems with bounded reconstruction error. The study introduces a novel “design-to-reality error decomposition theorem,” which decouples total error into five independently controllable components, enabling cross-modal composition of high-dimensional (3D–5D) primitives. Evaluated across six real-world data modalities, the method achieves expert-level quality with 98.1 ± 4.2% fidelity and successfully produces ten novel imaging designs that surpass the capabilities of any single modality.