🤖 AI Summary
This study addresses the limitation of existing methods in explicitly modeling point correspondence reliability and pose informativeness when localizing event cameras within LiDAR maps. To this end, we propose the CELL framework, which enables end-to-end learning of correspondence confidence via a differentiable probabilistic PnP solver. Specifically, CELL introduces a log-partition function constraint on weights to mitigate depth bias, designs a partial-completion depth representation coupled with a decoupled training strategy to prevent gradient interference, and incorporates edge-matching refinement for enhanced accuracy. Extensive experiments demonstrate that CELL significantly outperforms the LEAR baseline on the M3ED and DSEC datasets, reducing translation and rotation errors by up to 26.9% and 15.8%, respectively.
📝 Abstract
Localizing an event camera against a pre-built LiDAR map can be cast as dense optical-flow estimation between a rendered depth view and an event image, followed by a Perspective-n-Point (PnP) solver over the induced 3D-2D correspondences. Existing pipelines rely on geometric consensus during pose estimation, but do not explicitly model the reliability or pose informativeness, i.e., how strongly a correspondence constrains the camera pose, of individual correspondences. We show that the natural way to learn it -- using the per-correspondence error to constrain the learning of confidence -- suffers from a depth-dependent bias: small pixel errors reside predominantly at large depths and do not lead to high pose informativeness. Instead, in our method (CELL), we learn a per-correspondence confidence end-to-end through the pose, using a differentiable probabilistic PnP whose log-partition term encourages weight configurations that yield a better-constrained pose distribution. The learned confidence is used in three ways: (i) it reweights the flow supervision in a decoupled training scheme that keeps pose gradients out of the flow/edge backbone; (ii) it drives a probabilistic correspondence selection at test time; and (iii) together with the network's edge-probability it weights a final edge-matching refinement. We further design a partial-completion depth representation that adds signal without hallucinating across large gaps. On M3ED and DSEC our full system improves over the LEAR baseline on the majority of the evaluated sequences: it reduces the median translation error by up to 26.9% and the median rotation error by up to 15.8%.