Vision-Wireless Fusion for Multi-User Localization: A Cross-Modal Transformer Approach

📅 2026-09-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出一种基于交叉模态变换器的视觉-无线融合框架,用于解决复杂城市环境中多用户定位问题,通过结合CSI和视觉信息提高定位精度。
📝 Abstract
Accurate multi-user localization is challenging in complex urban environments, where wireless measurements can become ambiguous under noise, blockage, and multipath, while visual observations provide complementary spatial context. This paper presents a vision-wireless fusion framework for multi-user localization using pilot-indexed channel state information (CSI). Orthogonal pilot indices preserve the identities of the communicating UEs in the CSI token sequence and localization outputs. The model encodes each pilot-indexed CSI observation as a query token and uses cross-attention to retrieve user-specific information from spatial visual memory. Self-attention among CSI tokens further captures inter-user interactions, while the resulting multimodal representations are used for user-wise localization. Experiments on different datasets show consistent improvements over model-based, CSI-only, and multimodal-fusion baselines. Further experiments evaluate the model under different wireless and visual conditions.
Problem

Research questions and friction points this paper is trying to address.

multi-user localization
urban environments
wireless measurements
visual observations
spatial context
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Modal Transformer
Pilot-Indexed CSI
Multi-User Localization
🔎 Similar Papers
No similar papers found.