🤖 AI Summary
This study addresses the Sim-to-Real transfer challenges in multimodal integrated sensing and communication (ISAC) models arising from reliance on labeled data and inconsistent simulation configurations. To this end, we propose a framework that automatically generates deployment configurations via natural language. Methodologically, a dual-agent architecture is constructed to coordinate scenario generation with task execution, incorporating structured domain knowledge to guide planning alongside verification-based feedback correction. Furthermore, agentic AI, mixture-of-experts (MoE), and few-shot learning are integrated to achieve dependency-aware planning. Experimental results on the DeepSense 6G dataset demonstrate that the proposed approach significantly improves vehicle detection accuracy, beam prediction performance, and planning correctness. This work provides an efficient paradigm for zero-shot cross-domain deployment of ISAC systems.
📝 Abstract
Multi-modal integrated sensing and communication (ISAC) enables environmental perception and reliable connectivity for intelligent wireless networks. Data-driven multi-modal ISAC models depend heavily on annotated real-world data to learn relationships across sensing and wireless observations, thereby constraining scalable deployment. Although synthetic data generation reduces the burden, adapting existing simulation pipelines to a target deployment requires consistent scene, sensing, wireless, and learning configurations, while mismatches among these coupled components impair sim-to-real transferability. To address the challenge, we propose an agentic artificial intelligence (AI) framework for sim-to-real multi-modal ISAC, named AIMS. Given a natural-language deployment request specifying the target task, deployment conditions, and real-data budget, AIMS derives a deployment-specific sim-to-real configuration and coordinates its execution to produce a deployment-specific task model. A two-agent architecture coordinates scene construction with task learning. A scene construction agent generates geographically grounded, synchronized sensing and wireless records from shared physical states, while a scene understanding agent configures task-relevant modalities and mixture-of-experts (MoE) learning for zero-shot inference or few-shot adaptation. Structured domain knowledge guides dependency-aware planning, while validation evidence supports feedback-driven revision of affected decisions. Experiments on the real-world DeepSense~6G dataset demonstrate improved vehicle detection and beam prediction over the considered simulation and fusion baselines. A separate orchestration benchmark evaluates task interpretation, dependency reasoning, and feedback-driven replanning across diverse deployment requests, showing improved plan correctness with structured domain knowledge and validation feedback.