Instruction-Conditioned Electromagnetic Spectrum Understanding via Budget-Adaptive Signal Tokenization

πŸ“… 2026-10-08
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the inherent trade-off between the prohibitive token overhead of dense encoding and the loss of local features under fixed compression when feeding raw I/Q signals into multimodal large models. To this end, we propose BATok, a budget-adaptive signal tokenizer that pioneers a budget-adaptive mechanism. It extracts features via lightweight multi-resolution branches and dynamically allocates token capacity by integrating local energy priors with learnable queries, projecting compact signal representations into the embedding space of vision-language models. Furthermore, we construct EMSpec-Instruct, the first multimodal dataset aligning raw signals, waterfall diagrams, and linguistic instructions. Experiments demonstrate that the proposed framework achieves superior performance across modulation recognition, structured detection, and language-conditioned signal localization, effectively balancing signal fidelity with computational efficiency.
πŸ“ Abstract
Electromagnetic spectrum monitoring increasingly requires flexible analysis beyond task-specific recognition and detection. Multimodal large language models offer a unified interface, but extending vision-language models (VLMs) to raw I/Q signals requires tokenization that balances fidelity against a strict budget. For signals, dense encoding causes token costs to grow with observation length, whereas fixed-resolution compression may discard short-duration or localized signal evidence. Thus, we propose \textbf{BATok}, a budget-adaptive signal tokenizer that adjusts token capacity to the input length while allocating that capacity according to the signal content. BATok constructs candidate representations from signal-derived features using lightweight multi-resolution branches, then combines a local energy prior with learnable queries to resample these representations into compact signal tokens. The number of tokens adapts to the input length while remaining strictly bounded. The resulting tokens are projected into the language embedding space of VLMs. We further introduce \textbf{EMSpec-Instruct}, a multimodal instruction dataset aligning raw I/Q signals, waterfall images, and language supervision for modulation recognition, structured detection, and language-conditioned signal grounding. Experiments show that BATok learns effective signal representations and achieves competitive performance across all tasks.
Problem

Research questions and friction points this paper is trying to address.

electromagnetic spectrum understanding
signal tokenization
budget-adaptive
multimodal large language models
I/Q signals
Innovation

Methods, ideas, or system contributions that make the work stand out.

Budget-Adaptive Tokenization
Signal Tokenizer
Multimodal Large Language Models
I/Q Signals
Instruction Dataset
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Lei Zhai
Lei Zhai
School of Artificial Intelligence, Xidian University, Xi’an, China
Z
Zhihao Chang
School of Artificial Intelligence, Xidian University, Xi’an, China
Shuyuan Yang
Shuyuan Yang
Xidian University
Professor
Z
Zhixi Feng
School of Artificial Intelligence, Xidian University, Xi’an, China