🤖 AI Summary
This study addresses the excessive token overhead of textual coordinate representations and the inherent trade-off between precision and range in fixed quantization for perception tasks in multimodal large language models (MLLMs). To overcome these limitations, this work proposes a dynamic vector decoding framework that pioneers the unification of diverse 2D and 3D perceptual representations into compact discrete tokens. Efficient integration is achieved through 1D vector sequence mapping, high-dimensional space discretization, and a lightweight de-tokenizer architecture. By transcending the constraints of traditional textual coordinates in unbounded spaces with high-precision requirements, the proposed approach demonstrates superior performance across multiple benchmarks. Furthermore, it significantly reduces token consumption and inference latency, establishing a versatile and efficient perception paradigm for MLLMs.
📝 Abstract
Multimodal large language models have made remarkable progress in bridging vision and language, facilitating various perception tasks essential for human-machine interaction, robotics, and autonomous driving. However, existing MLLM-based perception methods predominantly rely on text-based coordinate representation, which suffers from excessive token overhead, or fixed-range quantization, which suffers from range and precision constraints, especially for 3D domains with unbounded spatial range and high localization accuracy requirements. To address these challenges, we propose a dynamic vector decoding method named DVD, which unifies the representation of 2D and 3D perception tasks. Specifically, we first transform diverse perceptual representation (i.e., 2D bounding boxes, 2D masks, and 3D bounding boxes) into 1D vector sequences, which are then mapped to compact discrete tokens in the high-dimensional space. Then, a lightweight de-tokenizer enables seamless integration with MLLMs by decoding output tokens back to original 2D and 3D perceptual representations. Extensive experiments on 2D and 3D perception benchmarks including RefCOCO series, SUN-RGBD, KITTI, Hypersim, nuScenes demonstrate that DVD achieves superior performance in 2D and 3D tasks and reduces significantly the token overhead and inference latency. DVD provides an efficient and general framework for integrating perception capabilities into MLLMs, overcoming the inherent limitations of existing methods.