FloorSAV: Elucidating Spatial Audio-Visual Context with 2D Floormap for AV-LLMs

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of Audio-Visual Large Language Models (AV-LLMs) in 3D spatial reasoning caused by the absence of explicit global geometric mechanisms. To this end, we propose FloorSAV, a framework that innovatively introduces dynamic 2D maps as a synchronous stream to integrate multimodal cues—including 3D point clouds, camera trajectories, spatial audio, and semantic landmarks—and injects them into the large language model for joint spatial understanding within a single inference pass. Furthermore, we construct SAVED-Bench, a dedicated benchmark tailored for this task. Experimental results demonstrate that FloorSAV significantly enhances spatial reasoning performance on both SAVED-Bench and SAVVY-Bench, validating the substantial potential of incorporating precise geometric information for advancing multimodal spatial understanding.
📝 Abstract
While 3D spatial reasoning in dynamic egocentric environments is crucial for embodied intelligence, audio-visual large language models (AV-LLMs) lack explicit mechanisms to process and internalize global geometry directly from raw sensory streams. Existing approaches either require costly fine-tuning or underutilize the model's cross-modal reasoning capacities. In this paper, we propose FloorSAV, a novel framework that explicitly grounds spatial audio-visual context by rendering a dynamic 2D floormap. By integrating 3D point clouds, camera trajectories, spatial audio cues, and semantically grounded object landmarks, we inject this floormap into the AV-LLM as a synchronized stream with an egocentric video. AV-LLMs utilize their multi-modal capabilities to jointly reason over visual, auditory, and geometric cues in a single inference with floormap interpretation guidance. We further introduce SAVED-Bench (Spatial Audio-Visual Egocentric Benchmark with Dynamic Agents), constructing essential tasks of spatial capability in real-world scenarios: dynamic relativity, regional, and path reasoning QAs. FloorSAV improves AV-LLMs' spatial reasoning on various tasks from both SAVED-Bench and SAVVY-Bench. Studies with ground-truth floormaps demonstrate the substantial potential of FloorSAV with accurate spatial information.
Problem

Research questions and friction points this paper is trying to address.

spatial reasoning
audio-visual large language models
egocentric environments
embodied intelligence
cross-modal reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Audio-Visual Large Language Models
2D Floormap
Spatial Reasoning
Cross-modal Reasoning
Egocentric Benchmark
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.