MatchFusion: Explicit-Implicit Instance Matching for Spatio-Temporal Multimodal Autonomous Driving

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决多模态自动驾驶中实例匹配问题,提出MatchFusion方法,结合几何相似性和类别一致性初始化匹配,并利用实例嵌入优化关联,实现高效准确的信息交互。
📝 Abstract
Sparse instance representations provide a compact interface for spatial LiDAR-camera and temporal past-current interaction in multimodal perception and E2EAD. Effective interaction requires reliable instance correspondences despite geometric discrepancies and heterogeneous semantic representations. Attention-based methods exploit contextual semantics but often require specialized representation alignment, increasing computational overhead. In contrast, association based on structured object states is efficient and interpretable but lacks contextual evidence to resolve ambiguous matches. To combine these complementary strengths, we propose MatchFusion, a learnable instance matching and fusion module for spatio-temporal multimodal autonomous driving. MatchFusion initializes pairwise affinities using geometric similarity and category consistency, then selectively refines structurally plausible associations using instance embeddings. The resulting soft matchmap guides a common residual aggregation operator for adaptive information exchange. This unified matching-fusion formulation supports spatial LiDAR-camera and temporal past-current interaction, using multi-view image-plane geometry and motion-compensated BEV geometry as the respective structural priors. Experiments on nuScenes demonstrate consistent perception gains across diverse front-end configurations. Compared with a prior instance-centric fusion method, the MatchFusion-equipped system achieves higher perception accuracy while reducing FLOPs by 55.3% and GPU memory usage by 39.3%, with the matching-fusion module accounting for only 3.7% of total perception latency. Integrating temporal MatchFusion into SparseDrive further improves perception within an E2E framework without additional supervision. These results establish explicit-implicit matching as an effective and efficient mechanism for spatio-temporal instance interaction.
Problem

Research questions and friction points this paper is trying to address.

instance matching
multimodal autonomous driving
geometric discrepancies
heterogeneous semantic representations
spatio-temporal interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

explicit-implicit matching
spatio-temporal multimodal
instance matching
perception accuracy
computational efficiency
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xiaoyu Li
Harbin Institute of Technology, Harbin 150001, China
J
Jiajia Fu
Harbin Institute of Technology, Harbin 150001, China
L
Long Shi
Harbin Institute of Technology, Harbin 150001, China
Tianyu Du
Tianyu Du
Zhejiang University
AI SecurityAdversarial Machine Learning
R
Ruihang Li
Zhejiang University, Hangzhou 310027, China
X
Xian Wu
Harbin Institute of Technology, Harbin 150001, China
L
Lijun Zhao
Harbin Institute of Technology, Harbin 150001, China
Yingtao Zhang
Yingtao Zhang
Professor of Computer Science,Harbin Institute of Technology
Pattern RecognitionMachine LearningComputer VisionImage Processing
L
Lining Sun
Harbin Institute of Technology, Harbin 150001, China
R
Ruifeng Li
Harbin Institute of Technology, Harbin 150001, China